MODEL NOTES & AGENT SYSTEMS

AI systems you can evaluate

Source-aware notes on coding agents, evaluation, security boundaries, and provider changes—written for builders who need operational decisions, not model hype.

Browse all articles
START HERE

From AI claims to a system you can evaluate

Start with data boundaries, understand the runtime, then build an evaluation and a provider exit plan.

01

Define the data boundary

Separate retention promises, safety processing, and the controls your application still owns.

Read guide
02

Measure finished work

Build a task set and rubric that measure correctness, safety, cost, and variance.

Read guide
03

Prepare for provider change

Turn a platform change into concrete portability, fallback, and migration checks.

Read guide
THE LIBRARY

Continue exploring this topic

01

OpenAI Plans to End Cursor Model Access: What Developers Should Prepare Before November 12

OpenAI says it intends to wind down its Cursor model contract after SpaceX acquired the coding-tool company. Here is what is confirmed, what remains uncertain, and how teams can reduce model-provider lock-in now.

02

Tencent Hy4 Preview Explained: What 770B and a 1M Context Window Actually Change

Tencent has released and open-sourced Hy4 preview. Here is what the official release says about its 770B/49B MoE design, one-million-token context, access, pricing, and the evidence developers should still verify.

03

OpenAI’s Private Safety Processing: What Zero Data Retention Changes for Builders

OpenAI is previewing a safety layer for Zero Data Retention deployments. Here is what the announcement says, what it does not promise, and what teams should verify before relying on it.

04

Build a Repeatable LLM Agent Evaluation Before You Pick a Model

A practical evaluation design for comparing coding and tool-using models on the work your product actually needs.

05

Claude Text Watermarking Explained: Detection, Limits, and What It Means for Writers

Anthropic says future Claude models will use invisible statistical watermarks. The useful question is not whether a detector can label every paragraph, but what the signal can and cannot establish.

AI systems you can evaluate — aifincode