Context engineering: agent skills, MCP, and governed context
Prompt engineering was never enough for teams shipping software with models. Context engineering is the discipline of assembling what the model sees — instructions, tools, retrieved docs, repo layout, policies — under a fixed token budget, with versioning and review like any other interface.
Agent skills (packaged instructions + optional scripts) and MCP (Model Context Protocol servers that expose tools and resources to clients) are two concrete implementations of that idea. They are not magic; they are how you stop re-pasting the same ten paragraphs into every session.
Context is the real API surface
Every completion is f(context) → text | tool_calls. What goes into context is a design choice:
- System and developer messages (behavior, safety, output shape)
- Tool definitions and JSON schemas
- Retrieved passages (RAG) — untrusted data
- Selected files, diffs, logs — often trusted but still bounded
- Conversation history — summarized or windowed
Skills: operational playbooks, not “better prompts”
A skill in tools like Cursor is roughly: a named bundle of when to use, step-by-step procedure, constraints, and sometimes allowed commands or file patterns. Compared to a one-off prompt:
| Ad hoc prompt | Skill | |---------------|--------| | Lives in chat history | Lives in repo or org library, diffable | | Drifts per engineer | Reviewed in PR like runbooks | | Hard to discover | Invoked by name or auto-routed when relevant |
Skills shine for repeatable workflows: deploy checklists, security review rubrics, migration steps for a specific framework, “how we write ADRs in this repo.” They fail when used as a substitute for product requirements or automated tests — the model still hallucinates if the skill describes fantasy architecture.
How to write skills that hold up:
1. Scope narrowly — one skill per workflow, not “everything about backend.”
2. State prerequisites — required files, env vars, commands that must exist; say what to do when missing.
3. Prefer verification steps — “run npm test path”, “grep for X”, not “ensure quality.”
4. Version with the codebase — skill that references removed scripts is worse than no skill.
Skills are documentation the agent executes; keep them honest about what the repo actually contains.
MCP: tools and resources without bespoke integrations
MCP standardizes how a client discovers and calls tools and reads resources from a server (local or remote). Instead of N custom plugins per IDE, you implement one server per boundary: issue tracker, observability, internal docs, staging database read replica (read-only!), deployment API.
Design choices that matter in production MCP servers:
- Least privilege — separate servers per sensitivity; do not mount prod write tools next to doc search.
- Stable tool names and schemas — clients and skills depend on them; treat breaking changes like API v2.
- Timeouts and pagination — list_issues without limits will blow the context window.
- Auth — OAuth or short-lived tokens; never long-lived secrets in the client config committed to git.
Skills + MCP together
Typical split:
- Skill answers how we work here (branch naming, test commands, review checklist).
- MCP answers what the outside world looks like right now (open PR, trace id, ticket status).
Without skills, MCP dumps raw JSON into context and the model improvises process. Without MCP, skills go stale the moment Jira field names change unless someone updates the markdown.
Context budget and compression
Models have large windows; latency, cost, and distraction still punish you for stuffing everything. Patterns:
- Progressive disclosure — load skill summary first; pull full skill or file chunks on demand.
- Structured summaries — table of API errors, not full log tail, unless debugging that incident.
- Reference by pointer — “see
lib/retrieval.ts” with selective read, not whole repo @-mention by default. - Separate retrieval from reasoning — RAG or search MCP for facts; keep system instructions short and stable.
Governance in real orgs
What distinguishes mature adoption:
- Owners for each skill and MCP server (on-call when the agent misbehaves in that domain).
- PR review for skill changes that touch security, infra, or customer data handling.
- Allowlists of which MCP servers junior roles can enable on laptops with prod VPN.
- Audit of tool calls from automated agents (CI bots, cloud agents), not only human IDE sessions.
What this is not
- Not a replacement for types, tests, and code review.
- Not proof the model understands your domain — it follows instructions and statistics.
- Not exclusive to one vendor — the ideas port; MCP is an open protocol, skills are a packaging pattern.
Closing
Context engineering is how senior teams make LLM-assisted work repeatable and reviewable. Skills encode procedure; MCP encodes live integration; your codebase and CI still encode truth.
Start with one high-friction workflow you already document badly (release, incident, migration). Turn that into a skill, wire one read-only MCP source for facts, and measure whether the next run needed less rework — that is the only metric that matters.