---
title: Benchmarks
description: Two benchmarks that answer two different questions — real API savings, and mechanism correctness at scale.
sidebar:
  order: 7
---

Two benchmarks live in `benchmarks/` in the repo, and they're deliberately not the same kind of thing — don't read one as a stand-in for the other.

## `openrouter-savings` — real API calls, real token counts

Makes real calls to OpenRouter and records the real `total_tokens` OpenRouter returns. It's intentionally small — a handful of decisions, some repeated — because real calls cost money and hit rate limits. Every number it prints came from an actual API response.

```bash
OPENROUTER_API_KEY=sk-... pnpm --filter @runic-labs/benchmarks run start
OPENROUTER_API_KEY=sk-... pnpm --filter @runic-labs/benchmarks run start -- --json
```

Bump the repeat count to widen the sample without editing the file:

```bash
OPENROUTER_API_KEY=sk-... REPEATS_PER_DECISION=6 pnpm --filter @runic-labs/benchmarks run start -- --json
```

An actual run, unmodified:

```json
{
  "model": "openai/gpt-oss-20b:free",
  "decisions": 48,
  "uniqueDecisions": 8,
  "apiCalls": 8,
  "cacheHits": 40,
  "tokensSpent": 2582,
  "tokensSaved": 12910,
  "tokensWithoutRunic": 15492,
  "savingsPercent": 83
}
```

48 decisions resolved, only 8 real calls made, and 83% fewer tokens spent than resolving all 48 the normal way — measured from OpenRouter's own `usage.total_tokens`, not estimated.

:::tip
Never commit a real API key to the repo. Keep `OPENROUTER_API_KEY` as an environment variable only — it's read from `process.env` in `benchmarks/openrouter-savings/index.ts` and nowhere else.
:::

## `reuse-sweep` — synthetic, proves the mechanism at scale

Makes no network calls at all. It exists to answer a narrower, honest question: *given N decisions and a repeat rate of R%, does the cache reuse exactly the entries it should, and does the ledger's bookkeeping come out exactly right?* It uses a fixed, clearly-labeled synthetic cost per decision (320 tokens) instead of a real API response — this is a correctness proof, not a real-world savings estimate.

```bash
pnpm run benchmark            # table
pnpm run benchmark -- --json  # machine-readable
```

Override the sweep's shape without editing the file:

```bash
SWEEP_DECISIONS=10,50 SWEEP_REUSE_PCTS=0,50,90 pnpm run benchmark
```

An actual run:

```txt
Runic reuse-sweep (synthetic — proves the mechanism, not real-world savings)
Synthetic cost per unique decision: 320 tokens

 decisions  reuse%   unique     hits      spent      saved  savings%
---------- ------- -------- -------- ---------- ---------- ---------
        10       0       10        0       3200          0         0
        10      25        7        3       2240        960        30
        10      50        5        5       1600       1600        50
        10      75        2        8        640       2560        80
        10      90        1        9        320       2880        90
       100       0      100        0      32000          0         0
       100      25       75       25      24000       8000        25
       100      50       50       50      16000      16000        50
       100      75       25       75       8000      24000        75
       100      90       10       90       3200      28800        90
      1000       0     1000        0     320000          0         0
      1000      25      750      250     240000      80000        25
      1000      50      500      500     160000     160000        50
      1000      75      250      750      80000     240000        75
      1000      90      100      900      32000     288000        90
     10000       0    10000        0    3200000          0         0
     10000      25     7500     2500    2400000     800000        25
     10000      50     5000     5000    1600000    1600000        50
     10000      75     2500     7500     800000    2400000        75
     10000      90     1000     9000     320000    2880000        90
```

The `savings%` column tracking the `reuse%` column 1:1 at every scale is the actual point — it's confirming the cache does exactly what it claims, not marketing a number.

## Which one should I trust for my use case?

Neither, directly — your actual savings depend on how often *your* agent re-asks the same decision, which `reuse-sweep`'s reuse% column models but can't predict for you. Run `openrouter-savings` against your own real decision set if you want a number specific to your workload.

## Next

- [How it works](/how-it-works) — the signature and ledger mechanics both benchmarks exercise
- [Quickstart](/quickstart) — wire this into your own agent to get your own real numbers