Engineering notes and research.
I write about AI agent architectures, distributed systems, and the tools that connect them. Currently building at Microsoft.
The road ahead: What evidence would justify more autonomy?
Before removing an approval step, ask what evidence would justify the change, what the trial cannot show, and what would reverse the decision.
Recent Writing
The road ahead: What evidence would justify more autonomy?
Before removing an approval step, ask what evidence would justify the change, what the trial cannot show, and what would reverse the decision.
Incident resolution with AI: A hypothetical reference design
A fictional incident workflow follows an AI-assisted rollback from alert to verified recovery, with explicit authority, handoff, and failure constraints.
From vibe coding to verified: Practising independent diagnosis
A proposal for practising independent diagnosis in AI-assisted engineering, with synthetic exercises and tests of whether the learning transfers.
What sleeper agents mean for production AI systems
Sleeper-agent experiments show why cleaner evaluations do not prove harmful behaviour is gone, and why production checks need independent evidence.
Designing Decision Boundaries for AI Autonomy
How to turn autonomy policy into testable action conditions, using a hypothetical booking service to examine consent, reversibility, retries, and the cost of delay.
Five Ways to Question an AI Agent's Success
A proposed review framework for checking an agent's diagnosis, scope, lasting outcome, downstream effects, and context freshness beyond a successful tool call.
AI Governance Needs Shared Rules and Different Controls
How to keep shared AI governance rules while tailoring controls to each use, with explicit ownership, testable evidence, and reassessment when permissions change.
When Human Oversight Pays for Itself
A practical cost model for human review, with a hypothetical payback calculation, sensitivity checks, and cases where oversight does not break even.
Your AI Agent's Skills Might Be Malware
What the ToxicSkills audit found, what its numbers do not prove, and how to limit the damage from a compromised agent skill.
Graduated Automation: How to Scale AI Without Losing Control
How a procedure earns permission to act: scoped approvals, independent checks, and withdrawing autonomy when conditions change.
The Hidden Debt in AI-Generated Code
What public studies reveal about persistent code-quality issues, the limits of AI-versus-human comparisons, and reviewing architectural change.
AI and critical thinking: What the research actually shows
What an EEG experiment, a knowledge-worker survey, and a perspective paper tell us about AI, critical thinking, and meaningful human oversight.
The Amplification Thesis
Human alone and AI alone score about the same. Human plus AI wins. An animated case for amplification over replacement.
AI Comes for the Grunt Work, Not the Judgment
The accountant already answered what happens to a job when the machine arrives. A scrollytelling essay on AI, work, and the bubble question.
Why Human-in-the-Loop is Not Optional for Production AI Agents
The six-pillar argument for why HITL is not a transitional compromise but a structural requirement of any production AI system.
Decentralized Granular Access Control for Agentic AI Systems
Multi-layered access control architecture for agentic AI in critical cloud infrastructure.
The Case for Human-in-the-Loop: A Six-Pillar Framework
53 research papers argue HITL is a structural requirement of any production AI system.
Zero Trust in Azure: Strategies for Securing Cloud Infrastructure
Seven practical strategies covering NSGs, Private Links, Azure Policies, AVNM.
Autonomous Incident Resolution at Hyperscale
Multi-agent orchestration for autonomous network incident resolution with safety guarantees.
From Reactive to Autonomous: Evolution of AI Operations
Five-generation maturity model from manual troubleshooting to autonomous resolution.
MCP vs CLI: Comparative Analysis of AI Agent Tooling
Complementary transport layers evaluated across 14 dimensions with benchmarks.