Benchmark

Measured, not claimed

Three numbers appear on this page. Each comes from a test series with recorded conditions, and each carries those conditions with it. A fourth statement is deliberately missing: that the answers get better as a result. That has not been measured, so it does not appear here.

38.8%less prompt text

That much text the model no longer has to read at full trimming. In the cautious variant it is 14.4 percent.

Context Economy v1 · 50 runs · three scenarios · offline

19.9%less power per request

In 50 out of 50 runs less energy was used, without exception. Offset against the extra computation that trimming itself costs, 13.8 percent remain — 18.9 with the larger model.

MeluXina Energy v1 · a single A100 card · Luxembourg

0risk documents lost

In safe mode not a single document was lost that mattered for a sensitive question. Hit rate on high-risk cases: 100 percent.

Hale Model Bench v0 · three models · no high-risk errors

Less text, same substance

A token is a building block of text, roughly a syllable. Language models bill per token: less text means lower cost and less computing time.

What was measured is how much of the prompt can be cut without losing anything. At full trimming it is 38.8 percent, in the cautious variant 14.4.

The cautious variant only trims the context and leaves every instruction in place. It saves less, but it is the route where the least can go wrong.

Tokens across 50 runs

original token usage · 496,200

full prompt, no trimming · 9,924 per run

cautiously trimmed · 424,670

−14.4% · 8,493 per run

fully trimmed · 303,618

−38.8% · 6,072 per run

Less power

19.9%

less energy while reading the question. In 50 out of 50 runs, without exception.

Per million requests

without trimming163 kWh
with full trimming131 kWh
saved32 kWh

Trim the text, save the resources

That saved reading also saves energy is not self-evident — which is why it was measured separately, on a single graphics card at the MeluXina data centre in Luxembourg.

In all 50 runs consumption was lower. What is measured is the reading of the prompt, before the first word of the answer exists: 101 watt-seconds less per request, which is 19.9 percent.

Trimming itself costs computing time. Subtract it and 13.8 percent of savings remain with the small model, 18.9 with the large one — where the confirmed range lies between 16.1 and 21.6 percent. The advantage grows with model size because the extra computation stays equally expensive.

The kilowatt-hours above are extrapolated from the measurements, not measured: one value per request, times a million. The large model is an open 14B standing in for others, not a statement about any vendor.

What is not measured

Whether the answers become better through trimming is open. The test report says so itself: „A quality test of the trimmed prompts is entirely missing — this benchmark counts tokens, not answers.“

What is documented is something smaller that holds: nothing is lost. No risk document, no passage, no file that would have mattered for a sensitive question.

Where the numbers come from

SeriesWhat it showsMeasured
Context Economy v1How much prompt text can be saved. Offline, 50 runs, three scenarios.3 August 2026
MeluXina Energy v1Whether the saving also saves power. A100, pilot and main run.4 August 2026
Frontier Proxy v0What happens with larger models. 14B as a stand-in.8 August 2026
Hale Model Bench v0Whether routing questions is reliable. Three models, no high-risk errors.28 July 2026

The words used in the reports

WordWhat it means
TurnOne question-and-answer run. „50 turns“ means: asked and answered fifty times.
TokenA building block of text, roughly a syllable. The billing unit of language models.
PromptEverything the model reads before answering: question, instructions, documents.
PrefillThe reading of the prompt. This is where the measured power consumption occurs.
Watt-secondThe smallest energy unit used here. 101 Ws is about a 10-watt lamp for ten seconds — that much is saved per request.
Kilowatt-hour3.6 million watt-seconds. The unit power is billed in. That is why the large figures here are in kWh.
7B, 14BModel size in billions of parameters. Larger usually means cleverer and more expensive.
Risk documentA file that matters for a sensitive question.