Blog
LLMMCPSkillsArchitecture

Context engineering: agent skills, MCP, and governed context

·8 min read

Prompt engineering was never enough for teams shipping software with models. Context engineering is the discipline of assembling what the model sees — instructions, tools, retrieved docs, repo layout, policies — under a fixed token budget, with versioning and review like any other interface.

Agent skills (packaged instructions + optional scripts) and MCP (Model Context Protocol servers that expose tools and resources to clients) are two concrete implementations of that idea. They are not magic; they are how you stop re-pasting the same ten paragraphs into every session.

Context is the real API surface

Every completion is f(context) → text | tool_calls. What goes into context is a design choice:

  • System and developer messages (behavior, safety, output shape)
  • Tool definitions and JSON schemas
  • Retrieved passages (RAG) — untrusted data
  • Selected files, diffs, logs — often trusted but still bounded
  • Conversation history — summarized or windowed
Senior teams treat context assembly as code: deterministic ordering, explicit truncation policy, redaction of secrets, and metrics on token usage per feature flag.

Skills: operational playbooks, not “better prompts”

A skill in tools like Cursor is roughly: a named bundle of when to use, step-by-step procedure, constraints, and sometimes allowed commands or file patterns. Compared to a one-off prompt:

| Ad hoc prompt | Skill | |---------------|--------| | Lives in chat history | Lives in repo or org library, diffable | | Drifts per engineer | Reviewed in PR like runbooks | | Hard to discover | Invoked by name or auto-routed when relevant |

Skills shine for repeatable workflows: deploy checklists, security review rubrics, migration steps for a specific framework, “how we write ADRs in this repo.” They fail when used as a substitute for product requirements or automated tests — the model still hallucinates if the skill describes fantasy architecture.

How to write skills that hold up:

1. Scope narrowly — one skill per workflow, not “everything about backend.” 2. State prerequisites — required files, env vars, commands that must exist; say what to do when missing. 3. Prefer verification steps — “run npm test path”, “grep for X”, not “ensure quality.” 4. Version with the codebase — skill that references removed scripts is worse than no skill.

Skills are documentation the agent executes; keep them honest about what the repo actually contains.

MCP: tools and resources without bespoke integrations

MCP standardizes how a client discovers and calls tools and reads resources from a server (local or remote). Instead of N custom plugins per IDE, you implement one server per boundary: issue tracker, observability, internal docs, staging database read replica (read-only!), deployment API.

Design choices that matter in production MCP servers:

  • Least privilege — separate servers per sensitivity; do not mount prod write tools next to doc search.
  • Stable tool names and schemas — clients and skills depend on them; treat breaking changes like API v2.
  • Timeouts and pagination — list_issues without limits will blow the context window.
  • Auth — OAuth or short-lived tokens; never long-lived secrets in the client config committed to git.
MCP does not replace your backend. It is adapter glue between the agent runtime and systems engineers already operate.

Skills + MCP together

Typical split:

  • Skill answers how we work here (branch naming, test commands, review checklist).
  • MCP answers what the outside world looks like right now (open PR, trace id, ticket status).
The agent loop: skill selects procedure → MCP fetches live data → skill says how to interpret → model proposes patch → human or CI verifies.

Without skills, MCP dumps raw JSON into context and the model improvises process. Without MCP, skills go stale the moment Jira field names change unless someone updates the markdown.

Context budget and compression

Models have large windows; latency, cost, and distraction still punish you for stuffing everything. Patterns:

  • Progressive disclosure — load skill summary first; pull full skill or file chunks on demand.
  • Structured summaries — table of API errors, not full log tail, unless debugging that incident.
  • Reference by pointer — “see lib/retrieval.ts” with selective read, not whole repo @-mention by default.
  • Separate retrieval from reasoning — RAG or search MCP for facts; keep system instructions short and stable.
Measure tokens in / tokens out / tool round-trips per task type when you argue about which skill to trim.

Governance in real orgs

What distinguishes mature adoption:

  • Owners for each skill and MCP server (on-call when the agent misbehaves in that domain).
  • PR review for skill changes that touch security, infra, or customer data handling.
  • Allowlists of which MCP servers junior roles can enable on laptops with prod VPN.
  • Audit of tool calls from automated agents (CI bots, cloud agents), not only human IDE sessions.
“Everyone paste the company policy into ChatGPT” is the shadow IT you are replacing — skills and MCP are the sanctioned path if governance keeps pace.

What this is not

  • Not a replacement for types, tests, and code review.
  • Not proof the model understands your domain — it follows instructions and statistics.
  • Not exclusive to one vendor — the ideas port; MCP is an open protocol, skills are a packaging pattern.

Closing

Context engineering is how senior teams make LLM-assisted work repeatable and reviewable. Skills encode procedure; MCP encodes live integration; your codebase and CI still encode truth.

Start with one high-friction workflow you already document badly (release, incident, migration). Turn that into a skill, wire one read-only MCP source for facts, and measure whether the next run needed less rework — that is the only metric that matters.