Public evidence snapshot, generated deterministically from the manifest by the validator. A merged change is not presented as maintainership; an owner-controlled counter is not presented as adoption.
Electrical Engineering PhD researcher at CUHK. I maintain public tools for reproducible engineering, Python packaging diagnostics, local-first data work, and explainable security evidence.
I try to make each claim inspectable: public releases, runnable entry points, documented limits, and direct links to upstream work. A download counter or a self-submitted project listing is not presented as independent adoption.
The six repositories pinned above are the current featured set: BenchLineage, VulnFuse, WillItBreak, FrontierTrials, TermScope and ColdShelf.
| Project | What it does | Verifiable entry points |
|---|---|---|
| BenchLineage | Records experiment provenance, calibration, uncertainty and evidence, with portable ELN import/export. | PyPI · v0.3.7 · five accepted ELN contributions |
| VulnFuse | Correlates findings from SARIF, Trivy, Grype, Snyk, CycloneDX, OSV and CSV without hiding merge blockers or scanner disagreement. | browser workbench · v0.4.24 · security model |
| WillItBreak | Diffs a package's public API between two versions and reports only the breaking changes that reach your call sites, with file and line numbers. Zero dependencies. | README |
| FrontierTrials | Runs capability trials against frontier models behind one config and compares the runs, privately. | try it · study report |
| TermScope | Plots Arduino, ESP32 and STM32 telemetry in a terminal over serial, pipes or SSH, with CSV record/replay. | PyPI · v0.4.1 · hardware reports wanted |
| Project | What it does | Verifiable entry points |
|---|---|---|
| ColdShelf | Builds a private searchable catalog of unplugged drives, including snapshots, duplicate evidence and physical-location notes. | latest release · quick start · limitations |
| contextcost | Measures how much LLM context a repository costs to read, identifies the generated/vendored/data files that waste it, and re-measures after a proposed cut so the saving is real, not estimated. | PyPI · v0.5.0 · GitHub Action · MCP server |
Experimental projects — useful, but not yet stable or broadly validated.
| Project | What it does | Verifiable entry points |
|---|---|---|
| OhmJudge | Answer-free, auditable electrical-engineering model evaluations — no API key required. | README |
| DidYouLearn | Outcome-based evaluation for AI tutors — which one actually helps you understand. | README |
| EvalInt | Lint your LLM eval set: reliability, items scored against a reference. | docs |
| ResearchBench | A working researcher's running comparison of AI systems on real research tasks — no API keys, no synthetic datasets, judgment by the person who needed the answer. | design doc |
See all repositories and published Python packages. The projects above are the small set I currently use to represent my maintenance work.
Recent owned-project maintenance includes VulnFuse PR #73 (v0.4.24), TermScope PR #8 (v0.4.1), JSONXray PR #7 (v0.2.1), BenchLineage PR #20 (v0.3.5) and SlowImports PR #6 (v0.2.1). These owner-authored and owner-merged releases, and a 2026-08-15 snapshot recording 692 rolling-month BenchLineage download events, demonstrate active maintenance. CONTRIBUTIONS.md holds the full upstream log.
Three things, all publicly inspectable:
- Build tools I use. Reproducible-engineering, Python-packaging diagnostics, local-first data and explainable-security tooling, maintained openly with releases and runnable entry points (above).
- Fix things where I am a user. 21 accepted pull requests across 12 upstream projects — eLabFTW, TheELNFileFormat, SampleDB, Astropy, CycloneDX, Keycloak, Plotly.js, rclone, Syft, tox and regl-line2d. I reproduce each issue locally and ship a test with the fix.
- Help maintainers decide. Root-cause analyses on upstream issues and tested reviews of third-party changes in the same communities.
The full, dated log with per-PR detail is in CONTRIBUTIONS.md.
- I am the primary maintainer of the owned repositories featured above.
- External work is described as merged or open contributions unless a project publicly grants broader responsibility.
- Synthetic or demo data is labelled as such; it is never presented as a user study or production deployment.
- Security and research tools publish their assumptions and failure boundaries.
- Stars, downloads, users, benchmarks and testimonials are never fabricated.