You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Free Apache-2.0 skills for Claude Code, plus trigger-probe — the open-source tool we used to measure all 18 of our paid skills: 144/180 fires when it should, 180/180 stays quiet when it should, 8 failures published.
SFX Lead Intelligence Command Center: local-LLM hub plus lead dashboard, quality lifted 61 to 99 percent via a ground-truth eval harness. Showcase; source private.
RL-style eval measuring intent/action divergence in frontier agents: model acknowledges a correction, then acts on the stale value anyway. 3 scenarios, 655 trials on claude-haiku-4-5, Sonnet 4.6, GPT-5.4, and Gemini 3.1 Pro Preview.