PHD / RESEARCH CV · RELIABLE AGENT EVALUATION
Jiaming Wei
Research Engineer — Reliable Agent Evaluation & Adaptive Agentic Systems
I study how to evaluate agentic systems when the evidence itself is unreliable: task outcomes vary across reruns, automatic judges make systematic errors, and final success labels can hide failures in trajectories, tools, providers or evaluation infrastructure. My current research asks what evidence is reliable enough to support adaptive decisions about perception, models, tools, memory, retries and test-time computation.
Research & Publications
Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents
UCL MSc · 2026Jiaming Wei, Zekun Wu, Adriano Koshiyama, María Pérez-Ortiz
Accepted at EMNLP 2026 Workshop REALM
Submitted to NeurIPS 2026 Workshop VLM4RWD · under review · arXiv:2608.06171
- Controlled comparison of DOM / SoM / vision and phantom representations, preregistered before the main runs.
- Core result: a routing value–learnability gap — value does not imply learnability.
- Tracks task-level failures and representation differences, with activation patching / linear probes for diagnosis.
- Large reproducible research system with heterogeneous-compute orchestration, automatic recovery and 1K+ research tests.
Education
University College London (UCL), Department of Computer Science
Sep 2025–Sep 2026MSc Artificial Intelligence for Sustainable Development · Department of Computer Science
Programme completion: 21 Sep 2026 · dissertation result pending
Primary supervisor: María Pérez-Ortiz · Second supervisor: Zekun Wu
MSc dissertation: When Is Expensive Perception Worth Paying For? Measuring the Ceiling, the Predictability and the Economics of Representation Routing in Web Agents
Xi'an Jiaotong University
Sep 2021–Jul 2025BEng Automation · GPA 3.82 · weighted average 87.95
Research / Engineering Experience
Holistic AI
Research Intern · London · Jun–Sep 2026 · Completed- Owned an end-to-end red-team/evaluation measurement line and separated target behavior, provider-side filtering and grader outcomes instead of collapsing them into one ambiguous binary result.
- Audited judge/grader FP/FN and failure modes against independent labels, then integrated the measurement logic into the working platform.
- Later built an NYC Local Law 144 audit pipeline: deterministic Python decision logic, evidence/logging and reviewable reporting, validated on real data and demonstrated live at the end of the internship.
Selected Research Systems
redteam-under-test
PUBLIC · 2026Puts the target, execution harness, provider behavior, judge and independent gold labels on the same measurement chain. The system asks not only whether the target failed, but whether the evaluator itself is trustworthy.
- 166 plugins mapped to local generation mechanisms; 103 enabled after end-to-end verification.
- Target-conditioned attack generation; a zero-egress path can keep generation, execution and judging local.
- Found an upstream refusal shortcut that missed 8 of 16 breaches on one probe.
FinQA
POST-TRAINING ATTRIBUTION · 2026Controls protocol, answer-format SFT, explicit reasoning and RLVR separately to attribute score changes to capability, protocol alignment or reward exploitation.
- Switching back to the correct native chat protocol alone recovered roughly +20–30 points for some larger base models.
- 4B answer-only SFT showed no significant net gain; paired McNemar turned an apparent improvement into a testable claim.
- GRPO/RLVR exposed clear reward exploitation while task accuracy fell.
Research Methods & Engineering
LLM & Agent evaluation · benchmark design · red teaming · judge / grader calibration · failure taxonomy · paired tests / bootstrap · preregistration
Web / Computer-Use agents · VisualWebArena · DOM / SoM / vision representations · tool use · MCP · trajectory analysis · recovery
PyTorch · Hugging Face · TRL · LoRA / SFT · GRPO / RLVR analysis · protocol attribution · reward-failure analysis
Python · TypeScript · Postgres · Playwright · Docker · GitHub Actions · Linux · AWS / SST · structured logging · data provenance · reproducible harnesses
English / tests: TOEFL iBT 107 (R28/L28/S23/W28) · MyBest 109 · 21 Sep 2024 · GRE 321 excluding writing