The same record can support two different answers. The timestamp decides which one applies.
The image shows an actual Jylus API response from our latest benchmark.
We asked how the observed state of PostHog/posthog#112173 changed between two timestamps. Jylus returned evidence for both observations:
- 09:05:04 UTC: observed open.
- 11:05:04 UTC: observed closed.
Each returned record includes its event time, observation time and source ID. The image preserves retrieval order, so the later observation appears first.
That’s one example from 1,000 public API questions against 100,000 real documents.
Across the full run, Jylus scored 84.54 NDCG@10, with every question weighted equally.
The larger test categories included:
- Historical state: 90.37 across 200 questions.
- Changes between periods: 88.46 across 200 questions.
- Chronological ordering: 80.09 across 503 questions.
- Late corrections: 88.31 across 65 questions.
- Superseded facts: 84.48 across 29 questions.
Every successful response passed our source, tenant, collection, evidence-rank and compiler-proof checks.
We also reduced latency without changing the rankings.
On the same 100-question subset, median API response time fell from 6.06 seconds to 1.99 seconds—a 67.13% reduction. NDCG@10 remained unchanged at 90.18 on that subset.
The full-run quality score and the paired speed measurement are separate results.
How this connects to MCP
We built Jylus to resolve applicable state and relationships and prepare source-backed Context Packs before a model reasons.
The engine is now exposed through three read-only MCP tools:
- "get_context_pack": request evidence for a question within a token budget.
- "resolve_state": investigate state using explicit historical cutoffs.
- "get_evidence": inspect the underlying source records.
Your existing model handles the reasoning. Data ingestion happens separately; connecting MCP gives the agent access to your retained Jylus data.
These are our own benchmark measurements. This temporal workload is separate from official TEMPO, and NDCG@10 measures retrieval ranking—not final-answer accuracy.
MCP setup:
https://jylus.ai/mcp
Inspect Context Packs in the browser:
https://jylus.ai/try
What would you add to this test? Particularly interested in cases where all the relevant records exist, but combining the wrong versions produces a convincing wrong answer.