Morphing Agents and Computers - Building a State-of-the-Art Researcher
A snapshot of how GREP works today, why planning and execution matter more than people think, and how two founders built a system that now outperforms Google, Perplexity, ChatGPT, and others on major research benchmarks.

This piece is a snapshot of how Grep works today (mid-March 2026) and where we think this architecture goes next.
For historical context, we originally took everything we learned at Parcha and turned that into a harness for doing corporate due diligence. We wrote about that in Claude in a Box. Over time, we realized that the concept of due diligence for businesses was really a special case of something broader: there was a well-established industry forming around AI-assisted, in-depth research. So we evolved our harness into a general researcher that could operate across almost any domain.
If you want an agent to do its best work, you have to give it the right environment to work in. That means more than a list of tools or MCPs. It means giving the agent a full compute structure and the ability to take on the work required by the task in front of it, while also managing context efficiently at every step.
This industry is moving extremely quickly, and this system may look completely different in a couple of months. But we wanted to stop and document what we have, how we built it, and what we think comes next.
How GREP Works at 40,000 Feet
The easiest way to understand GREP is as one agent that morphs across stages of the research process - and when the work demands it, splits into parallel copies of itself.
GREP has multiple modes, and depending on the mode and the effort level, the shape of the system changes. Here I am describing the higher-effort mode: the one that can go deep for hours, doing as much research as needed.
This is not a multi-agent system in the typical sense. We do not have a zoo of specialized agents negotiating with each other. Instead, we have one agent that reshapes itself depending on the stage, and can spawn sub-agents that are clones of itself with narrower focus. Think of it less like a team and more like cell division - one researcher that multiplies when the problem requires it, then converges back into one voice.
The second most important point is that execution is not a single pass. It is a two-layered loop:
- Inner loop: The researcher spawns sub-agents, evaluates their results, identifies gaps, and spawns more sub-agents until it is satisfied with the coverage and quality of the research.
- Outer loop: After the report is written, reviewers and fact-checkers evaluate the entire output. If the report does not meet the quality bar - if claims are unsupported, sections are thin, or the research has gaps - the system goes back to the research phase and the inner loop runs again.
The agent decides when it is done. Not a timer, not a token budget. The agent reads the review, compares it against the adequacy criteria from the plan, and either ships the report or goes back to work.
The flow starts with the planner, which takes the user's question, any supporting information, files, and additional context. It synthesizes all of that and begins to shape the research process. That plan is both a proposal for how the research should be carried out and an opinionated view of how the system should operate.
The planner is a researcher too. It may need to gather context through web search before it can write a good plan - very much like a coding agent exploring a codebase before taking on a large implementation task. It explores the topic, asks the user clarifying questions if needed, and then writes an elaborated plan aligned to the capabilities of the expert that will execute it.
Planning
In difficult research, the quality of the plan often determines the quality of the outcome.
Planning matters because difficult research is not just a matter of searching more. It is a matter of spending the right amount of time understanding the question, clarifying the user's intent, identifying what is known and unknown, and deciding how the problem should be broken apart.
The planner is not just routing work. The planner is doing research. Careful planning requires context, and in some cases it requires interacting with the user to clarify the question, the user's intent, or the context behind the request. A good planner gathers enough information to shape the rest of the system without overcommitting too early.
Expert Shaping
One of the most important things the planner does is shape what the expert will become for the task ahead. We do not think of GREP as a set of separate agents negotiating with each other. We think of it as one super-agent that morphs over time. Planning is its own phase, with its own skills and objectives, because that phase determines what the later researcher will need to be.
If the task is about financial analysis, the planner has to identify that the researcher will need the skills, tools, and ability to work with financial datasets, run analysis, and rely on reliable and up-to-date data sources. That is especially important because language models do not know everything, and more importantly, they must know what they do not know. Planning is where we decide how the system will complement the model's general knowledge with web access, tools, command-line interfaces, and computation.
Skills, Tools, and What the Model Doesn't Know
Planning is also where we decide which skills from our broader skill library should be loaded into the expert without overwhelming the context window. The planner selects tools, MCPs, CLIs, or specialized capabilities that might help the researcher operate more effectively. That might mean specialized news and real-time search, or it might mean coding tools, data analysis tools, or something even more specific.
The output of planning is a detailed file, and in some cases additional files with supporting structure. That file shapes the rest of the run. It includes an opinionated view of how the research should proceed, which capabilities should be loaded, what the major tasks are, what good outcomes look like, and what the final review checklist should contain.
Execution
Once the plan exists, execution becomes a problem of task shaping, context preservation, and structured synthesis. This is where the two-layered loop comes to life.
The first thing the researcher does is load the skills, tools, MCPs, and CLIs prescribed by the planner. But the researcher still has agency. It can adapt the plan, reshape it, and make decisions about how to pursue the goal. The plan is guidance, not a hard-coded script.
The Inner Loop: Research Until Satisfied
The first major execution step is turning the plan into a set of tasks. We use tasks heavily, and they are extremely powerful because they allow us to define work in a way that is specific, independent where possible, and suitable for concurrency.
This is where our architecture begins to look map-reduce-like. If the planner identifies five subtopics that can be researched independently, the researcher can spawn five sub-agents to go pursue those lines of inquiry. If a task needs a different type of expert - like a data analyst - it can spawn one with a different configuration, tools, and objective.
The inner loop does not stop after one round of sub-agents. The orchestrating researcher evaluates the results, identifies gaps in coverage or quality, and decides whether to spawn more sub-agents, deepen existing lines of inquiry, or move on. This loop continues until the researcher is satisfied that it has enough material to produce a strong report.
This is not a fixed number of iterations. The agent reads the evidence, compares it against the objectives in the plan, and makes a judgment call. Some questions resolve in one round. Others require three or four rounds of increasingly targeted research.
Context is Precious
The inner loop works because the main agent is disciplined about preserving its own context window. The orchestrating researcher does not try to hold every detail of every branch of research in its active context. Instead, it coordinates, synthesizes, and writes the final research while sub-agents use their own context windows to do the deeper work.
The file system is what makes this practical. Sub-agents do not send all of their work back through one giant message. They write files. In fact, they often write multiple files: the long analysis itself, and then smaller supporting documents that point the main agent to the most important sections. That way the orchestrator can read the summary or pointer file, decide which sections matter, and only read the relevant portions of the longer analysis.
This is one of the central ideas in GREP: context is precious. Effective use of context matters more than simply having a large context window. You need to preserve attention for synthesis, judgment, and writing - not waste it on moving large quantities of intermediate text around the system.
Tasks also help with persistence and recovery over long-running jobs. When research runs for a long time, the agent may compact and lose some active memory of the objective or the plan. The plan is always available in the file system to be read again, but tasks and task dependencies also preserve structure. They act as a durable representation of what needs to happen, what has already happened, and what depends on what.
Writing the Report
Writing the final report is more complex than it looks. Very long outputs are hard for language models to produce reliably in one pass. So instead of asking the agent to write a hundred-thousand-character report from scratch in one shot, we use what we think of as a pixelated-to-picture-perfect process.
The writer starts by generating a table of contents into a file. Then it writes thesis sentences or early scaffolding for sections and subsections. Then it expands section by section, gradually refining the report into something complete and cohesive. We do not use multiple writers stitching whole sections together because that often reads like multiple authors pasted into one document. The main writing agent owns the voice and cohesion from start to finish.
The Outer Loop: Review and Iterate
After the report is written, reviewers and fact-checkers take over. The final checklist defined in planning is read by reviewer agents, which perform their own validation, support claims with evidence, and identify areas where more research or correction is needed. They write their reviews into files.
Here is where the outer loop kicks in. The main agent reads the reviews and makes a decision: is the report good enough, or does it need to go back to the research phase? If the reviewers flag unsupported claims, thin sections, or missing perspectives, the system does not just patch the text. It goes back into the inner loop - spawning new sub-agents, gathering more evidence, and synthesizing again before rewriting the affected sections.
The agent decides when it is done. It reads the review, compares it against the adequacy criteria defined during planning, and either ships the report or goes back to work. This is what makes the system fundamentally different from a single-pass pipeline: the quality bar is enforced by the agent itself, not by a predetermined number of steps.
Benchmarks
We did not build GREP to win benchmarks, but the benchmarks made it clear we had built something unusual. They give us a way to understand how far we have come on certain dimensions and how we compare against the very companies and teams defining the category.
DRACO
DRACO is Perplexity's benchmark. It measures research quality across 100 questions in 10 domains - law, finance, medicine, technology, and more. Each answer is evaluated on factual accuracy, breadth and depth, presentation, and citation quality.
Our high-effort agent scored 78.6%, outperforming every entrant. Perplexity's best configuration came in at 70.5% - on their own benchmark. That's an 8.1 percentage point lead.
The domain breakdown tells a more interesting story. Grep wins 9 of 10 domains. The only loss: Personalized Assistant, by 1.5 points.
The largest gaps are in areas that require careful research methodology: Needle in a Haystack (+12.4pp), UX Design (+14.8pp), Shopping/Product (+10.5pp), and General Knowledge (+10.1pp). These are exactly the kinds of tasks where having a structured planning and execution pipeline - rather than a single-pass search - makes the biggest difference.
Grep also leads on all 4 rubric axes: Factual Accuracy (+7.5pp), Breadth & Depth (+7.2pp), Presentation (+3.0pp), and Citation (+14.5pp). The citation gap is particularly notable - our architecture's emphasis on evidence tracking and source verification through file-based context pays off directly.
DeepSearchQA
DeepSearchQA is Google's benchmark. It evaluates multi-hop search problems where the system must follow several steps of reasoning and information gathering to arrive at a specific answer. 896 questions, each requiring factual correctness.
Grep scored 84.5% factual correctness. That puts us ahead of Perplexity (81.9%, the unofficial SOTA before us) and well ahead of Google's own Gemini Deep Research Agent (66.1%), which created the benchmark.
What makes that result more interesting is that we did not achieve it with our highest-effort agent. We achieved it with our medium-effort agent. In GREP, effort does not simply mean doing more work. It often means being more focused, more deliberate, and less excessive. In benchmarks like this, excessive retrieval or overly verbose answers can actually hurt performance. GREP's focus helped here.
The thing worth noting here is the floor, not the ceiling. The lowest-performing category, Finance & Economics, still hits 79.5% on 132 questions. The spread from best to worst is only about 13 points. That matters because it means GREP is not gaming any particular question type - the architecture generalizes. Most deep research systems have a spike in one domain and a crater in another. Ours stays flat because the planning step adapts the research strategy per question rather than relying on a fixed retrieval pattern.
DeepResearch Bench
DeepResearch Bench is an independent benchmark built by researchers at the University of Science and Technology of China and Metastone Technology. It includes 50 PhD-level questions in English and 50 in Chinese, evaluated with the RACE scoring framework across comprehensiveness, insight, instruction following, and readability.
Grep scored 56.27, edging out Cellcog Max (56.13) and a competitive field of research systems. The margins are tight at the top - 0.14 points separating first and second.
The most surprising detail: our reports were even higher quality on the Chinese portion (56.42) than on the English portion (56.13), despite the fact that we had spent effectively zero time optimizing specifically for Chinese.
That result says something important about the architecture. When you give a very strong base model the right structure, the right tools, and the right environment, you often do not need to over-index on narrow prompt tuning or specialized optimization for every domain. You give the agent expertise, you give it capabilities, and then you let it work.
Note: As of publication, the official leaderboard lists our previous harness version in second place (56.09). Evaluators are currently confirming our latest results (56.27), which are published in full at our benchmarks repo.
What's Next
The story of GREP so far is really the story of discovering what an expert needs in order to do excellent work.
Experts Need Skills
The first thing we learned is that an expert cannot exist without specialized capabilities. That means tools, skills, command-line interfaces, and increasingly the guided ability to write code. It is not enough to hand an agent a static list of APIs. You need to give it actionable capabilities that match the work in front of it.
Experts Need Computers
Giving an agent a computer turned out not to be a convenience but a requirement. It needs to write and run programs. It needs a place to hold data, previous research, code, notes, and intermediate outputs. The file system is not just storage - it is a practical extension of the model's working capacity.
One useful mental model is to think of the context window like L1 or L2 cache in a computer: fast, precious, and limited. The file system is more like a larger and slower memory layer. It is not as immediate, but it is much more durable and much more available. That distinction has become increasingly important in how we think about agent design.
The computer does not have to be one thing. For some customers, it is a very simple sandbox where untrusted code can be run safely. For others, it can be much more powerful. We already have customers exploring what happens when research agents have access to GPU clusters and can perform significantly more computationally expensive operations.
Experts Need Brains
We are increasingly realizing that an expert needs a brain. It needs to remember previous work, preserve useful information across runs, accumulate context over time, and eventually improve by learning from prior research. Language models by themselves do not do this well. They do not naturally become better at a job just because they have done it many times. So one of the areas we are most interested in now is building abstractions on top of them that let them remember and deepen their expertise over time.
General Tools Over Narrow Interfaces
On the skills side, we are also learning that it may be less important to keep adding handcrafted capabilities than it is to let the agent write code, execute code, and use general-purpose tools well. Unix, Bash, Python, package managers, data tooling, and the file system are often more powerful than an ever-growing library of narrow interfaces.
The theme, increasingly, is simple: give the agent guidance, give it agency, and let it cook.
A Personal Note
One of the most fun parts of this project is that it has largely been built by two people, using the system itself to run the company around it.
Parcha and GREP are down to two people - AJ and myself. Not exactly by design. Startups are very hard. But the two of us have been able to build all of this, run the business, move customers onto it, and continue innovating by using GREP itself.
We have GREP experts helping with go-to-market. We have GREP experts helping with finance. We have coding agents fixing bugs. We have a growing internal army of experts helping us research, operate, and build. In practice, it is the two of us working alongside an army of GREP experts.

That has made building the company this way incredibly fun. It also feels like a preview of what is coming next for companies more broadly: smaller teams, leaner organizations, and significantly more output per person.
If you want to learn how we are doing this, or if you want to partner with us to optimize how you and your company work using GREP, get in touch. We would be happy to share notes and give you access to some of the latest things we have been building.
We're just getting started.
Miguel Rios is co-founder and CTO of Grep. Previously Head of Platform Engineering at Brex and Head of Consumer Data Science at Twitter. @miguelrios