Parsica Recall
A fully deterministic benchmark. No LLM judge.
We don't post our numbers.
Benchmarks matter in their own way. But there are too many knobs: publish only the k@ that flatters, swap the judge model, redefine what counts as a correct answer. In practice a benchmark number is mostly a record of which knobs someone knew to tune. We run benchmarks constantly, as internal engineering instruments. We just don't treat the resulting numbers as evidence, ours included.
So there are no scores on this page, and there won't be.
More benchmarks than any one person should have to endure.
This posture wasn't theoretical. It came out of ablation week, and out of running more benchmark passes than any one person should have to endure. The full account of that stretch, and why it ended here, is what this page is for.
Deterministic. Configurable. Yours.
What we built instead: a fully deterministic benchmark, agnostic to the model behind it, constructed from your own corpus. You choose how many memories it draws from, and it tests recall against the memory you actually keep, not a synthetic set tuned for a chart. A full-stack mode follows.
It's a fair test because you run it yourself, on your own data. There is nothing to game, because the test is your own history.
Did your agent recall the right memories, at the right time, when it needed to?
