Research Engineer — Reliable Agent Evaluation & Adaptive Agentic Systems
What evidence is reliable enough to support an adaptive agent decision? I study that question across web agents, evaluator reliability, trajectories and post-training.
I study how to evaluate agentic systems when the evidence itself is unreliable: task outcomes vary across reruns, automatic judges make systematic errors, and final success labels can hide failures in trajectories, tools, providers or evaluation infrastructure.
Reliable Agent Evaluation for Adaptive Agentic Systems
What minimum evidence is required before an adaptive agent is justified in changing its behaviour or learning from feedback?
My current direction treats evaluator error and outcome instability as part of the decision problem itself: when should an agent acquire richer perception, switch models, call a verifier, retry, ask a human, spend more test-time compute or abstain?
MSc Artificial Intelligence for Sustainable Development · 2025–2026
Research on representation choice and routing for Web / Computer-Use agents: first measure whether value exists, then ask whether that value is actually learnable.
“Routing Is Least Learnable Where It Is Most Valuable” was accepted at the EMNLP 2026 Workshop REALM.
The work combines preregistration, controlled representation comparisons, task-level failure analysis, representation probes and a recoverable heterogeneous-compute experiment system.
Web AgentsComputer UseEvaluationRepresentation RoutingResearch Systems
Holistic AI
Research Intern · London · Jun–Sep 2026 · Completed
Treated evaluation as a system rather than a score: from attack generation, target execution and grader auditing through to an executable audit pipeline and live delivery.
Owned an end-to-end red-team/evaluation measurement line and separated target behavior, provider-side filtering and grader outcomes instead of collapsing them into one ambiguous binary result.
Audited judge/grader FP/FN and failure modes against independent labels, then integrated the measurement logic into the working platform.
Later built an NYC Local Law 144 audit pipeline: deterministic Python decision logic, evidence/logging and reviewable reporting, validated on real data and demonstrated live at the end of the internship.
What did the model actually do: succeed, refuse, or cross the boundary?
behavior · task success
02→
Execution / Harness
Did the model cause the result, or did observation, tools, provider filters, scaffolding or environment change it?
trajectory · environment
03→
Judge / Grader
Can the evaluator itself be trusted? Rubrics, thresholds, shortcuts and FP/FN can change the conclusion.
FP / FN · calibration
04
Gold / Evidence
What ultimately anchors the claim: independent labels, statistical tests, provenance and replayable evidence?
labels · provenance
Web-Agent routing, Holistic, redteam-under-test, FinQA and the Model Observatory look like different projects, but they ask the same question: what does this result mean, and can the measurement chain that produced it be trusted?
paper-deslop2026A LaTeX-aware academic rewrite pipeline where models propose diffs and deterministic invariant gates protect numbers, citations, equations and terms.
Encode Persona2026Compares natural language, JSON/YAML, semantic labels and opaque-tag persona representations with ablations, cross-model probes and exact provenance.
Spatial Copilot2026Turns first-person narration into a persistent topological map, treating graph topology as truth and geometry as a visual hypothesis.
FitnessOS2026An offline-first SwiftUI fitness system with SQLite as source of truth, HealthKit, deterministic scheduling and a transactional outbox.
I want to make evaluation reliability part of the agent decision problem, not an afterthought.
Current Fall 2027 direction: reliable agent evaluation, evaluator-aware adaptation, trajectory-grounded supervision and resource-aware verification.