Inside the Agent
from first tool call to autonomous agent
120 hands-on articles across 13 modules — build real AI agents inside Cursor and Claude Code, the tools professionals actually use. Tool use, RAG, MCP, context engineering, memory, the Claude Agent SDK, multi-agent orchestration, evals, security, and production deployment — every lesson read, built, and shipped in a real editor, ending in benchmarked capstones.
Curriculum
13 modules · 120 topicsEvery article runs concept → build → ship — the idea, a hands-on lab in Cursor or Claude Code, and a graded task you keep. Where a lesson needs a prerequisite it links to Getting Started or Beyond the Prompt instead of repeating it.
- 14.01What This Path Is — and What It Assumes
- 14.02Installing and Touring Cursor
- 14.03Installing Claude Code
- 14.04Cursor's Modes and Inline Edits
- 14.05@-Context and Codebase Indexing
- 14.06Claude Code as an Agentic Terminal
- 14.07The Coding-Agent Landscape (2026)
- 14.08Safety-First Defaults
- 15.01Prompting vs Agentic Workflows
- 15.02The Augmented LLM
- 15.03What Is an Agent, Really?
- 15.04Anatomy of an Agent and the Think → Act → Observe Loop
- 15.05When NOT to Build an Agent
- 15.06Autonomy as a Spectrum
- 15.07Your First Agent, Felt Not Coded
- 16.01Chain-of-Thought, Revisited for Agents
- 16.02ReAct: Reason + Act
- 16.03Plan-and-Execute
- 16.04Reflection and Self-Critique
- 16.05The Evaluate-Improve Loop
- 16.06Reasoning Models and Test-Time Compute
- 16.07Agent Skills, Roles, and Capabilities
- 16.08Spec-Driven Development
- 16.09Decision Journaling
- 17.01How Tool and Function Calling Actually Works
- 17.02Designing Good Tools
- 17.03CodeAct vs JSON Tool Calls
- 17.04Calling External APIs from an Agent
- 17.05Querying Databases from an Agent
- 17.06Structured Outputs and Schema Validation
- 17.07Embeddings, Vector Stores, and Why RAG
- 17.08Building a RAG Pipeline in Cursor
- 17.09Agentic RAG
- 17.10Evaluating Retrieval
- 18.01Why MCP Exists
- 18.02MCP Architecture
- 18.03The Three Primitives: Tools, Resources, Prompts
- 18.04Adding MCP Servers to Cursor
- 18.05Adding MCP Servers to Claude Code
- 18.06Using Prebuilt Servers
- 18.07Building Your First MCP Server with FastMCP
- 18.08Building an MCP Client
- 18.09Deploying a Remote MCP Server
- 18.10MCP Authorization: OAuth 2.1 and PKCE
- 18.11Code Execution with MCP
- 18.12MCP Security and the Trust Boundary
- 19.01The Statelessness Problem
- 19.02Context Rot and "Lost in the Middle"
- 19.03Compaction and Summarization
- 19.04Structured Note-Taking and Tool-Result Pruning
- 19.05Just-in-Time vs Pre-Loaded Context
- 19.06Memory via Files: CLAUDE.md
- 19.07The CLAUDE.md Hierarchy and @-Imports
- 19.08Cursor Rules and AGENTS.md
- 19.09Progress Files and Shift Logs
- 19.10Database and Vector-Store Memory
- 19.11Memory Research: MemGPT/Letta and Sleep-Time Compute
- 20.01The Extensibility Stack, Mapped
- 20.02Slash Commands
- 20.03Skills: Authoring and Auto-Triggering
- 20.04Subagents
- 20.05Hooks: Deterministic Control
- 20.06The Plugin System and Marketplaces
- 20.07Essential Plugins, Toured
- 20.08Plugin Safety
- 20.09Building and Publishing Your Own Plugin
- 20.10Plan Mode and Multi-Context-Window Workflows
- 20.11The Initializer / Harness Pattern
- 20.12Failure Modes, Fixes, and Cursor Parallels
- 21.01Agent SDK vs the Messages API
- 21.02Install and First Query
- 21.03The Agent Loop Under the Hood
- 21.04Configuring the Agent
- 21.05Choosing Your Model
- 21.06Built-in Tools
- 21.07Custom Tools and In-Process MCP
- 21.08Streaming and Sessions
- 21.09Checkpointing and Permissions
- 21.10Cost, Caching, and Reliability Engineering
- 21.11Harness Thinking and the No-Code Counterpoint
- 22.01Multi-Step Planning
- 22.02The Five Workflow Patterns
- 22.03Orchestration vs Choreography
- 22.04When Multi-Agent Breaks
- 22.05Subagent Orchestration in Claude Code
- 22.06LangGraph
- 22.07CrewAI and AutoGen
- 22.08Choosing Your Orchestration Approach
- 23.01Why Evals First
- 23.02Objective vs Subjective Evals
- 23.03Response, Step, and Trajectory Evaluation
- 23.04pass@k vs pass^k
- 23.05Error Analysis with Traces
- 23.06Eval Platforms and Benchmarks
- 23.07The Epistemics of Benchmarks
- 23.08Guardrails and Bounded Autonomy
- 23.09Deterministic Enforcement
- 23.10Prompt Injection and the Lethal Trifecta
- 23.11Human-in-the-Loop
- 23.12Responsible AI in Practice
- 24.01From Notebook to Service
- 24.02Containerizing Agents with Docker
- 24.03CI/CD for Agents
- 24.04Observability and Tracing
- 24.05Cost and Latency Management
- 24.06State and Checkpointing at Scale
- 24.07Hosting the Agent SDK Securely
- 24.08Scaling, Enterprise Integration, and AgentOps Maturity
- 25.01MCP Governance and the Registry
- 25.02AGENTS.md as a Cross-Tool Standard
- 26.01Watching the Agentic Web: A2A and Agent Commerce
- 26.02Computer-Use and Browser Agents
- 26.03Multimodal Agents
- 26.04Human-Agent Collaboration
- 26.05Application Patterns Across Verticals
- 26.06Production Case Studies and the ROI Reality
- 26.07Capstone A: The Deep-Research Agent
- 26.08Capstone B: The OAuth-Secured MCP Platform
- 26.09Capstone C: The Autonomous Feature Builder
- 26.10Capstone D: Your Own Agent and Portfolio