An independent experiment suggests that language models can substantially reduce metered output when instructed to rewrite completed information in the compressed style once used for costly telegrams. The Telegraph Test project reports reductions of roughly 40% to 49% for several tested systems, while other models were able to recover the information from those shortened records at about the same rate as from ordinary prose.

The technique, called cablese or telegraphese, removes articles and filler, uses abbreviations and retains facts, numbers and proper nouns. It differs from a codebook, which substitutes predefined words for whole phrases. In the reported tests, codebook-style replacement produced savings of about 10%, whereas asking models to use a concise telegraph register yielded much larger reductions on providers' own token meters. Lowercase output mattered because all-capital formatting increased token use.

The benchmark used 50 passages and about 1,300 questions, according to the project account. Cross-family tests asked models to answer questions using records compressed by a different model. The reported recovery ratios ranged from 0.99 to 1.10 compared with answers based on plain text, meaning no tested comparison in that matrix favored the uncompressed version. The author has published a harness, frozen ledgers and a notebook intended to make the results reproducible.

Placement in a workflow was crucial. The experiment found worse results when a model was told to reason and compose directly in cablese. Performance improved when ordinary content was settled first and then compressed into a record for another machine to read. The proposed uses are therefore agent memory, scratch records and system-to-system handoffs, not user-facing writing or a substitute for normal reasoning.

The approach also failed economically for at least one tested model. The report says GPT-5 mini could not disable reasoning and spent much more effort producing the compressed version, causing billed output to rise despite shorter visible text. That result underlines the need to measure total provider charges rather than assume fewer displayed words always mean a lower bill.

There are important limitations. The work is an independently published benchmark, not a peer-reviewed finding, and the author says a strict control using a generic instruction such as “write as tersely as possible” has not been run. Such a comparison would help determine whether the historical telegraph framing adds anything beyond ordinary concision. Results may also change across providers, model versions, languages and types of material.

For developers, the practical takeaway is narrow but testable: after information has been finalized, a controllable model may be able to create a cheaper machine-readable record without materially reducing retrieval accuracy. Any production use should validate both comprehension and the full token bill on the exact models involved.