Xinyang (Alan) Chen

Agentic workflows, harness tooling, prompt evaluation — measured, not asserted.

[work]

claude-config

The infrastructure and the measurement discipline behind the rest: a Rust harness, MITM capture of live agent traffic, a replacement system prompt, and a runner-neutral corpus for evaluating prompts.

Not packaged for anyone else’s use — it is the instrument, not a product.

cv-final-project

The same workflow applied to real research: autonomous agent loops on a SLURM GPU cluster, for 6.8300. The contribution is the loop design, the concurrency isolation and the verification — the loops did the work.

[background]

MIT. Systems and compilers, ML and robotics, competitive programming (IOI gold, USAMO). Reproducible environments — Nix, Docker, a self-hosted homelab — because inspecting agent behaviour at the API level needs somewhere to stand.