- Total
- 1,935 ms
- First token
- 283 ms
- Cost
- $0.000098
Fastest + cheapest
Agent Routing Lab
Nebius Token Factory reference implementation
ONE CONTROLLED RUN · REAL API EVIDENCE
A working Token Factory + MCP comparison with preserved latency, cost, caching, and recovery evidence—and a follow-up review of an unreliable quality judgment.
One recorded observation. The evaluator’s preference is not a verified quality win. This public view is key-free and cannot spend tokens.
OBSERVED RESULT
Fastest + cheapest
Evaluator pick
EVIDENCE LEDGER
Load the sample to inspect the interface, or connect locally for a live run.
10 September 2026The evaluator missed a factual error. Route B says system-message or memory features do not count toward token limits. A system message still consumes input tokens. Memory depends on what is retrieved into context; it is not a blanket exemption.
B also assumes that the full conversation history is sent on each call. The supplied context does not establish that. The evaluator nevertheless awarded B 5/5 for groundedness and preferred it.
Revised interpretation: the capture proves that the model preferred B, not that B is more accurate. No quality winner is established. The original responses, scores, timings, costs, and exported trace are unchanged.
The next evaluation should check individual claims against source evidence, use an independent review rubric, and repeat across cases. A separate controlled experiment is needed to establish the effect of context size on latency.
The walkthrough predates this review and shows the original interpretation.
The prompt describes a suspected context slowdown. This trace does not measure a before/after context latency baseline; that requires a separate repeated experiment with only context size changed.
This surface will show two model routes, the MCP boundary, streaming first-token time, returned token usage, calculated cost, retry history, and one deliberately limited evaluation.
LangGraph → Tavily → Token Factory
A separate three-node LangGraph example performs agentic search with the official Tavily SDK, selects evidence, then passes bounded sources to a Token Factory-compatible answer step. One keyless Tavily search was captured live; its final summary was deliberately local, so no additional model tokens were spent.