No test cases yet.
Things you've told it about your pool, so you don't have to say them twice. It never stores your test results — those live in PoolMath and it reads them fresh every time. Forget anything here and it stops using it.
Loading…
--gen-context, which writes one situating sentence per passage — has no
button here and should not get one. Its output is sidecar files under
skill/references/.context/ that belong in git: both retrieval
legs read them, so they have to ship in the image, which means they have to be committed. A
server generating them would be doing ~1,800 model calls of work that nobody could commit and
that the next deploy would throw away. Run it from a checkout:
hsecret run OPENROUTER_API_KEY -- dotnet run --project src/Assistant.Evals -- --gen-context,
review the diff, commit the sidecars.SKILL.md) and can’t be edited here — changing it is a code change and a deploy, so the eval suite always grades exactly what members are served. Shown verbatim below.