Skip to content

Commit bd6fa5b

Browse files
committed
docs(site): disclose that the streaming mock's first token is instant
An audit reconciling our LiteLLM added-TTFT (~40 ms p99) against BerriAI's own bench (~0 ms) found the gap is by construction: our mock emits the first token instantly, so TTFT includes the gateway's first-write network behaviour (e.g. a server that never sets TCP_NODELAY eats a ~40 ms Linux delayed-ACK stall on the first content frame); a mock that waits before its first token, or that times the first SSE event of any kind, reports near-zero by construction. Disclose this so the number is not misread as "40x slower than LiteLLM claims". The mechanism-specific per-gateway note waits on an on-rig A/B confirmation.
1 parent dff172d commit bd6fa5b

1 file changed

Lines changed: 1 addition & 0 deletions

File tree

‎site/index.html‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -129,6 +129,7 @@ <h3><a id="lnk-memory" target="_blank" rel="noopener">memory/run.sh</a> Memory</
129129
<h3><a id="lnk-stream" target="_blank" rel="noopener">stream/run.sh</a> Streaming</h3>
130130
<ul>
131131
<li>The mock answers stream:true with a paced SSE stream; the suite measures what the gateway adds on top of that pace.</li>
132+
<li>The mock's <em>first</em> token is instant (it then paces later tokens), so the first-token figure deliberately includes the gateway's first-write network behaviour, not just its compute. Against a real model, whose own first token takes hundreds of milliseconds, that component is largely hidden; here it is exposed on purpose. A benchmark whose mock waits before its first token, or that times the first SSE event of any kind rather than the first content token, will report a near-zero first-token overhead by construction.</li>
132133
<li><b>Added wait for the first token (p99)</b>, also called TTFT (time to first token): gateway first-content-frame time minus direct-to-mock, at concurrency 1.</li>
133134
<li><b>Added gap between tokens (p99)</b>: gateway content-frame gap minus direct-to-mock gap.</li>
134135
<li><b>Streams sustained</b>: max concurrent streams where at least 99.9 percent of expected frames deliver, no stream stalls beyond twice the pacing interval, and the stream error rate stays under 0.1 percent.</li>

0 commit comments

Comments
 (0)