From eliasstravik/skills (MIT, "eval-built agent workflows")
The closest public analog to our eval framework. Its most transferable ideas, prioritized:
1. Skill-activation (trigger) evals — trigger-eval.json (highest leverage)
A flat JSON array of {"query": "...", "should_trigger": true|false} that tests whether a skill activates on the right requests and stays silent on adjacent ones — independent of behavior. For a 60+ skill marketplace, routing is our real failure mode (copywriting vs copy-editing vs ad-creative vs emails; seo-audit vs ai-seo vs schema; pricing vs paywalls). We have no eval for it today.
- Seed 60/40 positive/negative, with adjacent-skill negatives (e.g. "write ad copy" should NOT fire
copywriting — it should fire ad-creative; "is my pricing page AI-readable" fires pricing, not ai-seo).
- Elias runs a trigger ratchet that optimizes the skill's
description against the query set and applies best_description verbatim. Our per-skill description triggers are exactly what this would harden.
2. The "eval-built" prove-need lifecycle (for NEW skills)
"A skill is not written, it is earned." Before building a skill, run its scenarios against a no-skill baseline; record verbatim failure evidence; saturate (stop when a fresh round adds zero new failure classes). If the baseline doesn't fail, don't ship the skill. Then every skill line traces to a real failure via a failure→assertion traceability table, and the skill starts as a bare core (H1 + one primitive, ≤20 lines) with details earned by observed failures. Strong editorial filter against bloat for a growing library.
3. Harness upgrades
The obra complement (completes #479)
Elias saturates the NO-skill baseline (proves the skill is needed). obra/superpowers pressure-tests the WITH-skill artifact (proves it's robust). They're two halves — Elias has no adversarial "try to break the installed skill" round, which is exactly the pressure-test-evals gap already noted on #479. Together = the full picture: prove-need → build → prove-robust.
Credit / license
eliasstravik/skills is MIT — distill-with-credit. Adopt the formats and lifecycle, credit the source. Companion to #479 (framework), #480 (benchmark), #484 (schema convergence).
From eliasstravik/skills (MIT, "eval-built agent workflows")
The closest public analog to our eval framework. Its most transferable ideas, prioritized:
1. Skill-activation (trigger) evals —
trigger-eval.json(highest leverage)A flat JSON array of
{"query": "...", "should_trigger": true|false}that tests whether a skill activates on the right requests and stays silent on adjacent ones — independent of behavior. For a 60+ skill marketplace, routing is our real failure mode (copywriting vs copy-editing vs ad-creative vs emails; seo-audit vs ai-seo vs schema; pricing vs paywalls). We have no eval for it today.copywriting— it should firead-creative; "is my pricing page AI-readable" firespricing, notai-seo).descriptionagainst the query set and appliesbest_descriptionverbatim. Our per-skilldescriptiontriggers are exactly what this would harden.2. The "eval-built" prove-need lifecycle (for NEW skills)
"A skill is not written, it is earned." Before building a skill, run its scenarios against a no-skill baseline; record verbatim failure evidence; saturate (stop when a fresh round adds zero new failure classes). If the baseline doesn't fail, don't ship the skill. Then every skill line traces to a real failure via a failure→assertion traceability table, and the skill starts as a bare core (H1 + one primitive, ≤20 lines) with details earned by observed failures. Strong editorial filter against bloat for a growing library.
3. Harness upgrades
failure-classes.json+grading-schema.json(expectations[]/ typedclaims[]/pass_rate/eval_feedback) — converge these with Converge eval rubric schema with Magister's eval runner (3-layer rubrics, brand fixtures, config-as-candidate) #484 (the Magister schema work).The obra complement (completes #479)
Elias saturates the NO-skill baseline (proves the skill is needed). obra/superpowers pressure-tests the WITH-skill artifact (proves it's robust). They're two halves — Elias has no adversarial "try to break the installed skill" round, which is exactly the pressure-test-evals gap already noted on #479. Together = the full picture: prove-need → build → prove-robust.
Credit / license
eliasstravik/skillsis MIT — distill-with-credit. Adopt the formats and lifecycle, credit the source. Companion to #479 (framework), #480 (benchmark), #484 (schema convergence).