8000
Skip to content
#

swe-bench

Here are 195 public repositories matching this topic...

coder_eval

Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.

  • Updated Aug 24, 2026
  • Python
Xenon

Xenon — Agent Harness for AI coding agents. 7 种可替换推理范式 + Evidence Runtime 验证闭环(工具结果是 Evidence,LLM 输出只是 Claim)+ 执行隔离边界 + 可复现评测链路。交互与 SWE-bench 评测共享同一组约束,故 40.0% 实例级通过率(同模型 A/B +6.7pp)对日常使用同样成立。Python,可当 CLI 用也可当库嵌入。

  • Updated Aug 23, 2026
  • Python
why-was-fable-banned

Fable-style spec + evidence gate for Claude Code + Codex. Makes Opus/Codex work under Fable-like discipline: blocks every edit until a deterministic spec passes, and there is no "done" without live acceptance evidence. Spec-first, verification-gated, forbidden-paths enforced.

  • Updated Jun 30, 2026
  • Python

An 18 notebook course that isolates and measures each component of agentic loop engineering on real, industry standard software datasets.

  • Updated Jun 29, 2026
  • Jupyter Notebook

Improve this page

Add a description, image, and links to the swe-bench topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the swe-bench topic, visit your repo's landing page and select "manage topics."

Learn more

0