Anthropic's engineering essay on agent architectures — workflows versus agents, five composable patterns, and why simple beats clever.
Topic
Agent reliability & evals
Keeping agents alive in production — observability, evals, and the debugging practice behind marketing systems that run unattended.
Editions
Every edition that led with this, newest first.
CrewAI 1.15.23 lets your agent pipelines be scored in production and survive rate limits — tracing-based evaluation scores each run, and throttled model calls now retry automatically. If agents carry campaign work, upgrade in staging first.
Read editionYour LLM observability will break if you upgrade LiteLLM this week — the new release candidate migrates its Langfuse tracing callback to SDK v4, forcing anyone using Langfuse to rewrite integration code before deploying.
Read editionAnthropic released a new model, claude-opus-5-5, callable directly from its Python SDK — if you route marketing tasks across different AI models, re-run your tests before switching anything in production.
Read editionAgents can now work from your company's scattered files and cite their sources — V7's new platform turns internal docs into agent-ready memory, replacing workflows that today depend on someone digging through shared drives.
Read editionOpenAI is opening an ad surface inside agentic experiences — Sponsored Agents lets you run paid campaigns inside agent answers, managed through HubSpot and Shopify. That is a new paid channel every marketing engineer must evaluate now.
Read editionpromptfoo shipped version 0.123.0 — the open-source eval harness now covers the latest frontier models, but a default switch for newer GPT models breaks some existing test configs. If you run model evals, test before upgrading.
Read editionn8n, the workflow automation tool many marketing teams use to wire LLMs into campaigns, shipped version 2.39.0 with fixes to its AI agent nodes and security hardening across webhooks and secrets. If you run production marketing automations through n8n, this is a release worth testing before the next campaign cycle.
Read editionn8n shipped three releases in quick succession, fixing the reliability and integration issues that break AI agent workflows in production — Anthropic thread recovery, task runner timeouts, webhook handling, and secrets failover. If you run marketing automation through n8n, this week's upgrade is a maintenance task with real payoff.
Read editionGoogle shipped stateless updates to MCP that let agents scale horizontally without sticky session state — if you're pushing marketing agents toward production, this removes a core infrastructure bottleneck that has made concurrent personalization and lifecycle workflows hard to parallelize.
Read editionGoogle DeepMind shipped a faster Gemini Flash variant and agentic video understanding — two releases that could change how you route model calls and automate creative QA. PostHog open-sourced agent skill templates and Langfuse added eval alerting for production LLM workflows.
Read editionThree platform ships this week lower the barrier for deploying AI agents in marketing workflows — Google's MCP stateless scaling, Vercel's Claude Managed Agents in Chat SDK, and Vercel's dashboard agent deployment. But Twilio's data showing 78% of consumers actively bypass AI agents is a reminder that infrastructure readiness and consumer willingness are different problems.
Read editionDatabricks introduced a database architecture pattern for high-throughput agent workflows on Postgres — if your real-time personalization depends on agent loops, traditional transactional databases may be your throughput bottleneck. Three other releases this week touch LLM eval workflows, marketing-to-deployment pipelines, and HubSpot's platform direction.
Read editionLangfuse shipped two releases this week upgrading LLM evaluation workflows and agent observability. If you trace marketing agent calls or run evals on campaign prompts, reusable eval filters, execution SLO metrics, and OpenAI-compatible agent support are the headline features worth pulling into your stack.
Read editionLangfuse shipped a rebuilt eval experience, LangChain patched a Claude agent crash mode, and Supabase plus Bedrock added infrastructure for governed, context-aware agents — all relevant if you ship AI in marketing systems.
Read editionLangfuse shipped agent turn tracing for multi-step approval workflows, OpenAI's SDK added Bedrock and shell streaming support, and a compact multilingual embedding model landed — three builder tools worth evaluating this week.
Read editionVercel shipped a one-command setup for coding agents that routes through its AI Gateway — if your team experiments with AI-assisted content workflows, this removes the infrastructure configuration overhead.
Read editionVercel shipped a one-command setup for coding agents that routes through its AI Gateway — meaning marketing ops teams can spin up AI-assisted landing page builds and campaign experiments without wrestling with complex config.
Read editionLangfuse shipped per-conversation tool approval for in-app agents — if you're building agent workflows, this reduces friction by letting users approve a tool once per conversation instead of every call. Also relevant: n8n hardened AI agent execution with better schema handling and leak prevention.
Read editionGoogle shipped a production-ready agent evaluation service in Gemini Enterprise this week — the first turnkey option for teams that need to measure agent quality without building custom testing infrastructure from scratch.
Read edition
Resources
The catalog entries that go with it.
Hundreds of importable marketing automation templates — lead capture, content generation, social posting, and CRM sync flows, many agent-driven.
Langfuse
Open sourceOpen-source LLM observability — traces, evals, and prompt management for the agents you run in production.
promptfoo
Open sourceOpen-source testing and eval harness for LLM apps — write test cases, run them across models, and catch regressions before they ship.
Frontier Brief
Get the next brief
What shipped, what matters, and what to try Monday. Written for marketing engineers.