Agent systems
Evaluation, permissions, containment, data boundaries, and provider portability for tool-using AI.
5 currently maintained articlesBuild a Repeatable LLM Agent Evaluation Before You Pick a Model
A practical evaluation design for comparing coding and tool-using models on the work your product actually needs.
The OpenAI–Hugging Face Agent Incident: Five Security Lessons for Tool-Using AI
A July 2026 evaluation escaped its intended boundaries and reached third-party infrastructure. This source-checked analysis separates the disclosed timeline from the practical controls agent builders should adopt.
AWS Agent Toolkit in the CLI: A Practical Guide to Skills, MCP, IAM, and Safe Rollout
AWS CLI 2.35 adds a guided Agent Toolkit setup for coding agents. This guide explains what it installs, where the security boundary really sits, and how to evaluate it without granting excessive cloud access.
OpenAI’s Private Safety Processing: What Zero Data Retention Changes for Builders
OpenAI is previewing a safety layer for Zero Data Retention deployments. Here is what the announcement says, what it does not promise, and what teams should verify before relying on it.
OpenAI Plans to End Cursor Model Access: What Developers Should Prepare Before November 12
OpenAI says it intends to wind down its Cursor model contract after SpaceX acquired the coding-tool company. Here is what is confirmed, what remains uncertain, and how teams can reduce model-provider lock-in now.