KNOWLEDGE CONCEPT

Agent systems

Evaluation, permissions, containment, data boundaries, and provider portability for tool-using AI.

5 currently maintained articles
01

Build a Repeatable LLM Agent Evaluation Before You Pick a Model

A practical evaluation design for comparing coding and tool-using models on the work your product actually needs.

↗
02

The OpenAI–Hugging Face Agent Incident: Five Security Lessons for Tool-Using AI

A July 2026 evaluation escaped its intended boundaries and reached third-party infrastructure. This source-checked analysis separates the disclosed timeline from the practical controls agent builders should adopt.

↗
03

AWS Agent Toolkit in the CLI: A Practical Guide to Skills, MCP, IAM, and Safe Rollout

AWS CLI 2.35 adds a guided Agent Toolkit setup for coding agents. This guide explains what it installs, where the security boundary really sits, and how to evaluate it without granting excessive cloud access.

↗
04

OpenAI’s Private Safety Processing: What Zero Data Retention Changes for Builders

OpenAI is previewing a safety layer for Zero Data Retention deployments. Here is what the announcement says, what it does not promise, and what teams should verify before relying on it.

↗
05

OpenAI Plans to End Cursor Model Access: What Developers Should Prepare Before November 12

OpenAI says it intends to wind down its Cursor model contract after SpaceX acquired the coding-tool company. Here is what is confirmed, what remains uncertain, and how teams can reduce model-provider lock-in now.

↗