Loading…
Specshift · Run Eval
Drop a docs URL. We score retrieval, agent task completion, structure, and drift against methodology v1.0 and render the scorecard — overall score, per-suite breakdown, every test — inline in the response. No public report URL or README badge yet; that's INIT-PRAG-004. Optional email gets you the result delivery when the queue is busy.
Methodology v1.0. Engine + methodology versions are pinned in every scorecard so re-running the same target produces a deterministic score. See /specshift/methodology for the rules.