PromptsRush
Prompts

Browse

All PromptsThe full curated libraryPrompts GalleryVisual, Pinterest-style browsingImage PromptsMidjourney, DALL·E & SDXLVideo PromptsRunway, Kling & SoraText & TemplatesChatGPT & Claude system prompts

Discover

CategoriesExplore prompts by topicAI ModelsSpecs, benchmarks & best promptsCompare ModelsNewSpecs, pricing & benchmarks side by sidePrompt PacksCommunity, passcode-protectedSubmit a PromptShare with the community

For Creators

Turn prompts into followers

Share passcode-protected prompt packs and grow your audience with Auto DM.

Start sharing
Marketplace

Explore

Shared PromptsPasscode-protected prompt packsAI SkillsNewInstallable Agent SkillsDesign SystemsNewLive themes & design tokens

Contribute

Submit a PromptPublish a prompt packSubmit a SkillShip an Agent SkillSubmit a DesignShare a design system

New · Skills

Teach your AI new tricks

Install ready-made skills for Claude, ChatGPT, Gemini, n8n & more.

Browse skills
Learn

Learning Tracks

Prompt EngineeringWrite prompts that deliverAI SkillsBuild & ship Agent SkillsAI AutomationWorkflows, agents & MCPDesign SystemsOn-brand UI with AI

More

Learning HubAll tracks · 40+ lessonsBlogGuides, news & deep diveseBooksPremium prompt packs & guides

100% Free

Learn AI, the practical way

From fundamentals to advanced across four hands-on tracks — no fluff.

Explore the hub
BlogApp
LoginSign Up
PromptsRush

The ultimate directory for finding, sharing, and managing production-ready AI prompts, system instructions, and advanced templates.

TwitterGitHubYouTubeInstagramEmail

Platform

  • Home
  • Browse Prompts
  • Android App
  • Marketplace
  • Skills
  • Categories
  • Submit a Skill

Top Categories

  • Image PromptPopular
  • Video Prompts
  • Text Templates

Company

  • Privacy Policy
  • Terms of Service
  • Contact Us

Get the Android app

The whole prompt gallery, free on Google Play.

Install

Subscribe on YouTube

New AI prompt & skills tutorials every week.

Subscribe

© 2026 PromptsRush. Crafted with & Passion.

All systems operational
HomeBlogTutorials
Tutorials

What Is Context Engineering? The Next Step Beyond Prompt Engineering

Context engineering decides what an AI model sees at every step: instructions, memory, retrieved docs and tool results. Learn how it differs from prompt engineering, the 4 core techniques, 5 ways context fails, and a step-by-step workflow.

P
PromptsRushOctober 7, 2026
•24 min read3 views

Advertisement

What Is Context Engineering? The Next Step Beyond Prompt Engineering

Context engineering is the practice of deciding everything a model sees before it answers: the instructions, the conversation so far, the memories, the documents, the tool definitions, the tool results and the format it must reply in. Prompt engineering is about writing one good instruction. Context engineering is about building the system that assembles the right information, at the right moment, for every step an AI agent takes.

Short version: if you have watched an agent forget a decision it made twenty minutes ago, call the wrong tool, or confidently repeat its own mistake, the wording of your prompt was rarely the problem. The context was. This guide covers the definition, how it differs from prompt engineering, the seven parts of context, the four techniques practitioners use to manage it (write, select, compress, isolate), the five ways context fails, and a step-by-step workflow you can apply whether you are building a support bot or just trying to get better work out of a coding agent.

What Is Context Engineering?

The term went mainstream in June 2025, in two posts on X. On June 19, 2025, Shopify CEO Tobi Lütke wrote:

"I really like the term “context engineering” over prompt engineering. It describes the core skill better: the art of providing all the context for the task to be plausibly solvable by the LLM." (Tobi Lütke, June 19, 2025)

Six days later, on June 25, 2025, Andrej Karpathy added his "+1" with the line LangChain later quoted in its own context engineering guide: in every industrial-strength LLM app, context engineering is "the delicate art and science of filling the context window with just the right information for the next step." He also named the trade-off in one line: "Too little or of the wrong form and the LLM doesn't have the right context for optimal performance. Too much or too irrelevant and the LLM costs might go up and performance might come down."

Neither of them invented the practice. Cognition had already called context engineering "effectively the #1 job of engineers building AI agents" in its June 12, 2025 post, and LangChain's Harrison Chase published a formal definition on June 23: "building dynamic systems to provide the right information and tools in the right format such that the LLM can plausibly accomplish the task." On September 29, 2025, Anthropic followed with Effective context engineering for AI agents, describing it as "the set of strategies for curating and maintaining the optimal set of tokens (information) during LLM inference."

Our working definition: context engineering is designing the system that decides what goes into a model's context window at each step, and what stays out. The phrase "each step" matters. Anthropic's simple definition of an agent is "LLMs autonomously using tools in a loop", and every pass through that loop rebuilds the context. If you are still fuzzy on where agents end and reusable capabilities begin, our guide to AI skills vs AI agents covers that line.

Prompt Engineering vs Context Engineering

Anthropic calls context engineering "the natural progression of prompt engineering", and Harrison Chase argues that "prompt engineering is a subset of context engineering." Both framings hold. Prompt engineering is still the craft of writing clear instructions. Context engineering decides what else sits next to those instructions, where it comes from, and when it gets removed.

Prompt engineeringContext engineering
Unit of workOne instruction or templateEverything in the context window for this call
Core questionHow should I phrase this?What does the model need to see right now, and what should it not see?
When it happensOnce, before the callOn every turn of an agent loop
Where it livesA chat box or a system promptCode, files, memory stores, retrieval pipelines, tool definitions
Inputs managedRole, task, format, examplesInstructions, history, memories, retrieved documents, tools and their results, output schema
Typical failureVague or ambiguous wordingMissing, stale, conflicting or excessive information
Skills involvedClear writingWriting plus retrieval, memory design, summarisation, tool design and evaluation
Best forOne-shot tasks, chat, content draftsAgents, long-running tasks, multi-step workflows, production apps

Picture it this way: prompt engineering polishes one card, while context engineering decides the whole hand the model is dealt on every turn. For the related split between a one-off prompt and a packaged capability, see AI Skills vs Prompts: What's the Difference?

Split comparison: a single glowing prompt card on the left versus a layered stack of instructions, memory, documents and tools feeding the same AI core on the right

Why Context Engineering Matters Now

Two things changed at once: models became agents, and context windows became enormous. The second did not solve the first.

Bigger context windows did not fix the problem

Google's Gemini API docs note that many Gemini models accept 1 million or more tokens. Capacity is not the same as attention, though. The 2023 paper Lost in the Middle found a U-shaped curve: models used information best when it sat at the very start or end of the input, and worst when it sat in the middle. With the answer buried mid-context, GPT-3.5-Turbo sometimes scored below its own closed-book result of 56.1%, meaning the extra documents actively hurt.

Chroma's July 2025 report Context Rot tested 18 models, including GPT-4.1, Claude 4, Gemini 2.5 and Qwen3, and found that performance "consistently degrades with increasing input length", even on simple tasks. On a conversational memory benchmark, every model did significantly better with a focused prompt of about 300 tokens than with the full history of about 113,000 tokens. Anthropic explains why: transformers create n² pairwise relationships for n tokens, so every token draws on a finite "attention budget". Even Google's own docs advise that "if you don't need tokens to be passed to the model, it is best to avoid passing them."

Agents multiply the problem

An agent appends every action and every observation to its context, then reads the whole thing again on the next step. Manus reported that its average input-to-output token ratio is around 100:1, with a typical task needing around 50 tool calls. Anthropic measured that agents use about 4× more tokens than chat interactions, and multi-agent systems about 15× more. OpenAI's API docs are blunt about the cost side: requests are stateless, and even when you chain responses with previous_response_id, "all previous input tokens for responses in the chain are billed as input tokens." If you run coding agents on API billing, our breakdown of Claude Code API pricing and usage cost shows how quickly those tokens add up.

Honest answer: the bottleneck moved. Harrison Chase's diagnosis is that, more often than not and especially as models improve, mistakes happen because the model was not "passed the appropriate context", not because the model is incapable.

The 7 Components of Context

Context is everything the model sees before it generates a response. Philipp Schmid's breakdown lists seven parts, and they line up with the components Anthropic and LangChain describe. Each one is a lever you can tune, and each one competes for the same window.

Seven translucent glass layers (instructions, user input, history, long-term memory, retrieved documents, tools and output schema) stacking into a single context window in front of an AI core

1. System instructions

The standing rules: role, goals, constraints, examples. Anthropic recommends writing them at the "right altitude": specific enough to guide behaviour, not so rigid that you hardcode brittle if-else logic, and not so vague that the model has to guess. It also recommends a few "diverse, canonical examples" over a laundry list of edge cases.

2. User input

The immediate task or question. In a chat this is the latest message; in an agent it might be a ticket, a webhook payload or a scheduled job description. Google's long-context guidance adds a practical placement rule: performance is usually better when the question goes at the end of the prompt, after the material it refers to.

3. Short-term memory (conversation state)

The running history of this session: messages, tool calls and results. Each API request is stateless, so your application decides how much of that history to resend. This is where compaction and trimming live, and where most token bloat comes from.

4. Long-term memory

Facts that persist across sessions: user preferences, project decisions, past corrections. Claude Code's auto memory, for example, loads the first 200 lines or 25KB of a MEMORY.md index into every session and reads topic files only on demand. Its docs suggest that multi-step procedures belong in a skill rather than the always-loaded instruction file, which is exactly the job AI skills are designed for: procedural knowledge that loads only when it is relevant.

5. Retrieved knowledge (RAG)

Documents, database rows or web pages fetched for this specific question. Retrieval-augmented generation was introduced in a 2020 paper by Lewis et al. that paired a model's built-in knowledge with a searchable external index, which can be swapped to update what the system knows without retraining. In context engineering terms, RAG is one way to select context, not the whole discipline.

6. Tools and tool results

Tool definitions tell the model what it can do; tool results are what comes back, and they are often the largest thing in an agent's context. The Model Context Protocol, open-sourced by Anthropic on November 25, 2024 and described by its docs as "like a USB-C port for AI applications", made connecting tools trivial. That convenience has a context cost: every server you connect adds definitions the model must read, which is worth remembering when you browse lists like our best MCPs for generating images with Claude.

7. Structured output

The shape the answer must take, such as a JSON object with named fields. A schema is context too: it narrows what the model produces, and it makes the output safe to feed into the next step without another parsing pass.

The Four Core Techniques: Write, Select, Compress, Isolate

In July 2025, LangChain grouped the strategies used by popular agents into four buckets. Writing context means "saving it outside the context window". Selecting means "pulling it into the context window". Compressing means "retaining only the tokens required to perform a task". Isolating means "splitting it up". Almost every context engineering trick you will read about fits one of these four.

Four glass cards in a row showing the context engineering techniques write, select, compress and isolate, each with a simple icon

Write: save context outside the window

The agent writes notes, plans and progress to a file or memory store, then reads them back later. Anthropic's research agent saves its plan to memory first because, in its words, "if the context window exceeds 200,000 tokens it will be truncated". Manus keeps a todo.md and rewrites it as it works, which pushes the plan back to the end of the context where the model pays most attention. A minimal version for a long coding task looks like this:

# NOTES.md (agent-maintained)
Goal: migrate auth from sessions to JWT, no downtime
Decisions: keep /login route; refresh tokens in httpOnly cookie
Done: token service, middleware, 14/20 route handlers
Next: remaining 6 handlers in api/billing/
Open bug: logout does not revoke refresh token (see test_logout)

If you build agents on LangChain, its official LangChain Skills pack covers LangGraph persistence and memory patterns that implement exactly this.

Select: pull in only what this step needs

Selection is RAG, but also memory lookup, tool filtering and rule loading. Anthropic describes a "just in time" approach: keep lightweight references such as file paths, queries or links, and load the content with tools only when needed. Claude Code drops its CLAUDE.md in up front, then uses glob and grep to find files on demand. Cursor's rules do the same at the instruction level, with rules that always apply, apply to matching file paths, or load only when the agent judges them relevant.

Tool selection matters as much as document selection. Drew Breunig summarised research showing tool descriptions start to overlap and confuse a model past about 30 tools, and that retrieving a shortlist of relevant tools gave as much as 3× better tool selection accuracy. When the knowledge you need lives on the web, selection starts with getting clean, LLM-ready text instead of raw HTML; that is the job of web data APIs, and our Firecrawl pricing guide covers what that costs.

FirecrawlDeveloper Pick

Turn Any Website Into Clean LLM-Ready Data

Scrape, crawl, map and search the web through one API that returns clean markdown or structured JSON. 1,000 free credits every month, no card required.

Free / from $16 per month

Affiliate link · We may earn a commission

Try Firecrawl Free

Compress: keep only the tokens that matter

Compression means summarising or trimming what is already in the window. Anthropic's Claude Code compaction passes the history to the model to summarise, preserving "architectural decisions, unresolved bugs, and implementation details" while discarding redundant tool output, then continues with that summary plus the five most recently accessed files. The lightest version is tool result clearing: once a raw search result has been used, the agent rarely needs to see it again.

The risk is losing a detail that only matters later. Manus's answer is to make compression restorable: drop a web page's content but keep its URL, drop a file's contents but keep its path. Cognition went further and fine-tuned a smaller model just to compress agent history.

Pro tip: Tune a compaction prompt the way Anthropic recommends: first maximise recall so nothing important is dropped, then tighten precision by removing filler. A summary that loses one key decision costs more than one that keeps three redundant lines.

Isolate: split work across clean context windows

Isolation gives a sub-task its own fresh context. In Anthropic's research system, subagents explore in parallel, each with its own window, and return condensed summaries; Anthropic notes a subagent may burn tens of thousands of tokens but hand back "often 1,000-2,000 tokens". Its multi-agent setup outperformed a single Claude Opus 4 agent by 90.2% on an internal research eval, and Anthropic says multi-agent systems use about 15× more tokens than chat. Sandboxes and state objects isolate context too: a large file or dataset can live in a variable or on disk while only the relevant slice reaches the model.

Isolation has a real counter-argument. Cognition's position is that parallel subagents make conflicting decisions because they cannot see each other's work, so it defaults to a single-threaded agent and shares full traces. My take: isolate read-heavy exploration (research, search, codebase questions) and keep write-heavy decisions in one thread.

TechniqueWhat it doesConcrete exampleWatch out for
WriteSaves state outside the windowNOTES.md, todo lists, memory files, saved plansWriting unverified claims that later poison the context
SelectPulls in only what this step needsRAG, just-in-time file reads, tool shortlists, path-scoped rulesRetrieving near-misses that distract more than they help
CompressShrinks what is already thereCompaction, tool result clearing, trimming old turnsDropping a decision that matters ten steps later
IsolateSplits work across separate windowsSubagents, sandboxes, state fields hidden from the modelParallel agents making conflicting assumptions; token cost
Free on Telegram
Free on Telegram
Want more guides like this?
Join the PromptsRush Telegram channel for exclusive prompts and step-by-step AI guides, delivered straight to your phone.
Join on Telegram
Exclusive prompt drops
Step-by-step guides
Free · leave anytime

Five Ways Context Fails

Drew Breunig's June 2025 post How Long Contexts Fail named four failure modes that are now standard vocabulary, and Chroma's research added a fifth, broader one. Knowing which one you are looking at tells you which technique fixes it.

FailureWhat happensDocumented exampleUsual fix
Context poisoningA hallucination or error enters the context and keeps getting referencedGemini's Pokémon agent had its goals and summary "poisoned" with wrong game stateValidate before writing to memory; restart or quarantine the thread
Context distractionThe context grows so long the model leans on history instead of reasoningOnce its context grew well past 100k tokens, the same agent favoured repeating past actions over new plansCompress and summarise; start fresh with a handoff
Context confusionIrrelevant information or tools get used anywayA quantised Llama 3.1 8b failed a GeoEngine benchmark query when given 46 tools, but succeeded with 19Trim the tool loadout; prune retrieved content
Context clashParts of the context contradict each otherMicrosoft and Salesforce researchers spread prompts across multiple turns: scores fell 39% on average, and o3 dropped from 98.1 to 64.1Consolidate into one clear spec; remove superseded instructions
Context rotRecall and accuracy fall as input length grows, even on simple tasksAll 18 models in Chroma's tests degraded with longer inputsKeep context short; put key facts at the start or end

One nuance: not every error should be removed. Manus deliberately leaves failed actions and stack traces in context so the model does not repeat them. The distinction is between evidence (a tool call that failed, with its real error message) and invented facts (a hallucinated file path written into the plan). Keep the first, purge the second.

A Step-by-Step Context Engineering Workflow

You do not need a framework to start. This is the order we recommend, from cheapest change to most involved, whether you are wiring up your own agent or tuning a tool like Claude Code or Cursor.

Step 1: Define "done" before you touch the context

Anthropic's research team started with a set of about 20 queries representing real usage and found that small samples were enough to see the effect of changes early on. Write five to twenty realistic tasks and the result you expect for each. Without them, you cannot tell whether a context change helped.

Step 2: Audit what the model actually sees

Dump the full context of one real run, not the prompt you think you are sending. Claude Code's /context command shows which memory files loaded; most frameworks have a tracing view. Then sort every block into the seven components and ask what each one is earning.

Context Audit

Ready to use
You are reviewing the full context sent to an AI agent on one real task. I will paste it below.

Task the agent was doing: [DESCRIBE THE TASK]
What went wrong (if anything): [DESCRIBE THE FAILURE]

1. Split the context into these components and estimate the share of tokens each uses: system instructions, user input, conversation history, long-term memory, retrieved documents, tool definitions, tool results, output format.
2. For each component, list what the agent actually needed for this task and what was irrelevant, duplicated or outdated.
3. Flag any contradictions between parts of the context.
4. Flag any claim that looks like an earlier hallucination being repeated.
5. Recommend specific cuts, and say for each whether to delete it, summarise it, fetch it on demand instead, or move it to a separate sub-task.

Context:
[PASTE THE FULL CONTEXT]
Generate in Genspark

Step 3: Cut instructions to the right altitude

Anthropic's advice is to start with a minimal prompt on the best model available, then add instructions and examples only for failures you actually observe. Cursor's rule docs say the same in plainer words: keep rules under 500 lines, reference files instead of copying them, and add a rule only when you see the agent make the same mistake repeatedly.

Step 4: Trim the tool loadout

Anthropic's test is simple: if a human engineer cannot say which tool should be used in a given situation, the agent will not do better. Merge overlapping tools, write descriptions that state when not to use a tool, and return compact results rather than raw dumps. If you are building your own MCP server, the MCP Builder skill walks an agent through designing well-scoped tools with evals.

Step 5: Decide what loads up front and what is fetched on demand

Stable, always-relevant facts (build commands, house rules, the output schema) belong up front. Large or situational material (docs, past tickets, files) should be fetched by reference when needed. Keep a stable prefix: Manus points out that something as small as a timestamp at the top of a system prompt invalidates the cache for everything after it, raising cost and latency.

Step 6: Add a memory file and a handoff summary

For any task that spans hours or sessions, have the agent maintain a notes file and write a handoff summary before the context is compacted or a new session starts.

Handoff Summary Before Compaction

Ready to use
We are about to clear this conversation and continue the task in a fresh session. Write a handoff note that a new agent with no memory of this session can act on immediately.

Include, in this order:
1. Goal: the original objective in one or two sentences, in the user's words where possible.
2. Decisions made: every decision that constrains future work, with the reason.
3. Current state: what is finished, what is half-done, and the exact files, records or URLs involved.
4. Open problems: unresolved bugs or questions, with the exact error messages.
5. Next three actions, in order.
6. Do not repeat: approaches already tried that failed, and why.

Rules: keep facts that were verified by a tool result; mark anything unverified as UNVERIFIED. Leave out pleasantries, raw tool output and anything already finished and irrelevant to the next steps. Stay under [WORD LIMIT] words.
Generate in Genspark

Step 7: Isolate heavy exploration in subagents

Anthropic found that vague subagent instructions caused duplicated work and gaps, and that each subagent needs an objective, an output format, guidance on tools and sources, and clear task boundaries. Give subagents questions to answer, not decisions to make.

Subagent Brief

Ready to use
You are a research subagent with your own clean context. Your output goes back to a lead agent that has limited room, so be thorough in your search and brief in your answer.

Objective: [ONE SPECIFIC QUESTION TO ANSWER]
Why it matters: [ONE SENTENCE OF CONTEXT FROM THE LEAD AGENT]
In scope: [SOURCES, FOLDERS OR TOPICS TO COVER]
Out of scope: [WHAT OTHER SUBAGENTS ARE HANDLING]
Tools to prefer: [TOOLS] Start with broad searches, then narrow.
Stop when: you can answer with evidence, or after [N] searches.

Return only:
- Answer: 3-6 sentences.
- Evidence: up to 8 bullet points, each with its source URL or file path.
- Confidence: high, medium or low, with one reason.
- Gaps: what you could not verify.
Generate in Genspark

Step 8: Trace, measure and iterate

Rerun your test cases after each change and track two numbers: task success and tokens per task. LangChain's advice is to make sure you can see your agent's full inputs and token usage before you start optimising; context engineering without traces is guesswork.

Context Engineering Checklist

  • Every block in the context can justify its tokens for this specific step.
  • Instruction files (AGENTS.md, CLAUDE.md, Cursor rules) stay short; Claude Code's docs target under 200 lines per CLAUDE.md file.
  • No two instructions contradict each other, superseded instructions are deleted rather than overridden, and failures are fixed by supplying missing information rather than piling on more rules.
  • Tools have distinct jobs, clear descriptions and compact outputs; the agent sees only the tools this task needs, not every MCP server you have connected "just in case".
  • The question and the most important facts sit at the start or the end of the context, not buried in the middle.
  • Long tasks keep a notes file, and the agent writes a handoff summary before compaction.
  • Old raw tool results are cleared or summarised once used; compression keeps references so it can be undone.
  • Memory writes are validated so a hallucination cannot become a permanent "fact".
  • Read-heavy exploration runs in isolated subagents that return short, sourced summaries.
  • You have test cases and traces, and you measure tokens per task alongside success rate.

Real-World Examples of Context Engineering

Coding agents: AGENTS.md, CLAUDE.md and Cursor rules

Coding agents are where most people meet context engineering first, usually as a Markdown file in the repo. AGENTS.md describes itself as "a README for agents" and, as of October 2026, says it is used by over 60,000 open-source projects. It works across Cursor, Gemini CLI, Jules, Aider, Zed, Warp, Windsurf and others, the nearest AGENTS.md in the directory tree takes precedence, and the format is now stewarded by the Agentic AI Foundation under the Linux Foundation.

Cursor's docs state the reason these files exist: "Large language models don't retain memory between completions. Rules provide persistent, reusable context at the prompt level." Claude Code's docs say each session "begins with a fresh context window"; it loads CLAUDE.md at the start, can read AGENTS.md when there is no CLAUDE.md, and re-reads the project-root CLAUDE.md after /compact so standing instructions survive compression. That is write, select and compress, all in one Markdown file. The same logic applies to design context: our post on how design.md improves AI coding results shows what happens when an agent gets a design system instead of guessing one.

Draft a Lean AGENTS.md

Ready to use
Act as a senior engineer onboarding an AI coding agent to this repository. Using the files I paste below (README, package manifest, CI config and folder tree), draft an AGENTS.md under 120 lines.

Include only what an agent cannot reliably infer from the code:
- exact install, dev, test, lint and build commands
- project layout in 5-10 lines, pointing to key folders rather than describing every file
- conventions that differ from the defaults for this language and framework
- rules for tests, commits and pull requests
- things the agent must never do (generated folders, secrets, production data)
- known gotchas, with the file where each one bites

Leave out: generic style advice a linter already enforces, full dependency lists, and anything that duplicates the README.
End with a list of folders that deserve their own nested AGENTS.md.

Files:
[PASTE README, MANIFEST, CI CONFIG, FOLDER TREE]
Generate in Genspark

Customer support bots

A support bot is a good illustration of all seven components working together. The system instructions set tone, escalation rules and what the bot may never promise. The user input is the customer's message. Short-term memory is this conversation; long-term memory is the customer profile and past tickets, selected by customer ID rather than dumped in full. Retrieval pulls the two or three policy articles that match the product and region. Tools look up orders and create refunds, and the structured output returns the reply plus fields such as an escalation flag and a reason code.

The classic failure here is context clash: an outdated refund policy retrieved alongside the current one. The fix is a retrieval index that only holds the current version of each policy, with effective dates in the metadata, rather than an instruction asking the model to "prefer the newest document".

Research agents

Research is the textbook case for isolation. Anthropic's Research feature uses a lead agent that plans, spins up subagents that search in parallel with their own context windows, and passes the findings to a citation step at the end. Anthropic reported that token usage alone explained 80% of the performance variance on the BrowseComp evaluation, which is why spreading the work across separate windows paid off. Consumer research agents follow similar patterns; our Genspark review looks at one of them in day-to-day use.

Genspark

Try Genspark — the AI super-agent

Genspark researches, plans and acts across the web for you — multi-step agentic workflows in one prompt.

Affiliate link · We may earn a commission

Try Genspark Free

My Take: Is Context Engineering Just Prompt Engineering With a New Name?

Partly, and that is fine. The core skill of saying exactly what you mean to a model has not changed, and good prompts still sit at the heart of every context. What changed is the job around the prompt. Once a model runs in a loop, calls tools and works for hours, the hard questions are about information flow: what to load, what to fetch, what to forget and what to hand to a fresh window.

Our recommendation: start small. Write a short AGENTS.md or CLAUDE.md, trim your tool list, and add a handoff summary to any task that outlives one session. Those three changes fix most of the failures people blame on the model. Reach for subagents and custom compaction only when your traces show you need them. Anthropic's own advice, "do the simplest thing that works", is still the best rule in this field.

Try Genspark — the AI super-agent

Genspark researches, plans and acts across the web for you — multi-step agentic workflows in one prompt.

Try Genspark Free

Keep Reading

  • How to Write Better design.md Files: apply the same lean-context thinking to design instructions for coding agents.
  • How to Create Your First Claude Skill: package a repeatable procedure so it loads only when needed.
  • How to Use Claude Code for Free: put CLAUDE.md and AGENTS.md to work without a paid plan.
❓

Frequently Asked Questions

8 questions answered

Context engineering is deciding what information an AI model sees before it answers: its instructions, the conversation so far, memories, retrieved documents, tool definitions, tool results and the required output format. The goal is to give the model everything it needs for the current step and nothing that distracts it.
Prompt engineering focuses on writing a clear instruction for a single call. Context engineering manages the entire context window across many calls, including memory, retrieval, tools and history. Anthropic describes context engineering as the natural progression of prompt engineering, and LangChain's Harrison Chase calls prompt engineering a subset of it.
No single person coined it. The term spread widely after Shopify CEO Tobi Lütke posted about it on June 19, 2025 and Andrej Karpathy endorsed it on June 25, 2025. Cognition had already used it in a June 12, 2025 post, and LangChain published a formal definition on June 23, 2025.
No. Retrieval-augmented generation is one technique inside context engineering, used to select relevant documents for a question. Context engineering also covers system instructions, memory, tool selection, summarisation, compaction and splitting work across subagents.
No. Research such as Lost in the Middle (2023) and Chroma's Context Rot report (2025) found that models use long inputs unevenly and that performance tends to drop as input length grows. Larger windows also cost more per call. A focused context usually beats a full one.
Context rot is the decline in a model's ability to recall and use information as the number of tokens in its context grows. Anthropic uses the term in its context engineering guide, and Chroma's 2025 study observed it across all 18 models it tested, even on simple tasks.
They are LangChain's four groups of context engineering strategies. Write saves information outside the context window, such as a notes file. Select pulls only relevant information back in, such as RAG or a tool shortlist. Compress keeps only the tokens a task needs, such as summaries. Isolate splits work across separate context windows, such as subagents.
A little. Keeping one task per conversation, starting a fresh chat with a short summary when a thread gets long, and storing standing preferences in custom instructions or a project file are all simple context engineering. It matters most once you use agents, coding tools or long multi-step tasks.
Keep learning on Telegram
Keep learning on Telegram
Don’t miss the next guide
Exclusive prompts, practical AI guides and the best new tools — shared with our Telegram channel. Free, one tap to join.
Join on Telegram
Exclusive prompt drops
Step-by-step guides
Free · leave anytime
Back to Blog

Table of Contents

In this article

  • 1What Is Context Engineering?
  • 2Prompt Engineering vs Context Engineering
  • 3Why Context Engineering Matters Now
  • Bigger context windows did not fix the problem
  • Agents multiply the problem
  • 4The 7 Components of Context
  • 1. System instructions
  • 2. User input
  • 3. Short-term memory (conversation state)
  • 4. Long-term memory
  • 5. Retrieved knowledge (RAG)
  • 6. Tools and tool results
  • 7. Structured output
  • 5The Four Core Techniques: Write, Select, Compress, Isolate
  • Write: save context outside the window
  • Select: pull in only what this step needs
  • Compress: keep only the tokens that matter
  • Isolate: split work across clean context windows
  • 6Five Ways Context Fails
  • 7A Step-by-Step Context Engineering Workflow
  • Step 1: Define "done" before you touch the context
  • Step 2: Audit what the model actually sees
  • Step 3: Cut instructions to the right altitude
  • Step 4: Trim the tool loadout
  • Step 5: Decide what loads up front and what is fetched on demand
  • Step 6: Add a memory file and a handoff summary
  • Step 7: Isolate heavy exploration in subagents
  • Step 8: Trace, measure and iterate
  • 8Context Engineering Checklist
  • 9Real-World Examples of Context Engineering
  • Coding agents: AGENTS.md, CLAUDE.md and Cursor rules
  • Customer support bots
  • Research agents
  • 10My Take: Is Context Engineering Just Prompt Engineering With a New Name?
  • 11Keep Reading
Telegram channel
@promptsrush
Exclusive prompts, free
Prompt drops and step-by-step AI guides, straight to your phone.
Join on Telegram
Free · opens in the Telegram app

Android app

Prompts in your pocket

The whole gallery, customisable prompts and one-tap hand-off to your AI apps.

  • Swipe the full gallery
  • One-tap to ChatGPT & Gemini
  • Free · no account needed
Get it on Google Play

Recent Posts

How to Turn a Product Idea Into a Landing Page With AI

Oct 5 · 15 min

18 Prompts to Audit Your Landing Page With AI

Oct 3 · 13 min

10 Must-Have AI Tools for Marketers

Oct 3 · 12 min

ImagineArt Review (2026): Features, Pros & Cons

Sep 17 · 14 min

ImagineArt Pricing (2026): Plans, Credits, Usage

Sep 16 · 11 min

Category

Tutorials

Advertisement

You May Also Like

Tutorials

How to Turn a Product Idea Into a Landing Page With AI

Oct 515 min
Tutorials

18 Prompts to Audit Your Landing Page With AI

Oct 313 min
Tutorials

23+ Best 1980s, 85s, 90s AI Video Prompts

Sep 1421 min
More guides on Telegram
Exclusive prompts · 100% free
Join free