Agent Routing Lab

Nebius Token Factory reference implementation

Checking environment

EVIDENCE LEDGER

No comparison recorded

Load the sample to inspect the interface, or connect locally for a live run.

Fastest observed—
Cheapest observed—
Evaluator—
Observed cost—

Review of the saved capture

10 September 2026

The evaluator missed a factual error. Route B says system-message or memory features do not count toward token limits. A system message still consumes input tokens. Memory depends on what is retrieved into context; it is not a blanket exemption.

B also assumes that the full conversation history is sent on each call. The supplied context does not establish that. The evaluator nevertheless awarded B 5/5 for groundedness and preferred it.

Revised interpretation: the capture proves that the model preferred B, not that B is more accurate. No quality winner is established. The original responses, scores, timings, costs, and exported trace are unchanged.

The next evaluation should check individual claims against source evidence, use an independent review rubric, and repeat across cases. A separate controlled experiment is needed to establish the effect of context size on latency.

The walkthrough predates this review and shows the original interpretation.

WHAT THIS RUN MEASURES Two model routes for the same diagnostic prompt

The prompt describes a suspected context slowdown. This trace does not measure a before/after context latency baseline; that requires a separate repeated experiment with only context size changed.

00

This surface will show two model routes, the MCP boundary, streaming first-token time, returned token usage, calculated cost, retry history, and one deliberately limited evaluation.

Open-stack extension

LangGraph → Tavily → Token Factory

A separate three-node LangGraph example performs agentic search with the official Tavily SDK, selects evidence, then passes bounded sources to a Token Factory-compatible answer step. One keyless Tavily search was captured live; its final summary was deliberately local, so no additional model tokens were spent.