In my previous post on Prioritization, we built agents that decide what to do next when the to-do list is longer than the time available. That’s a pattern for execution: given a fixed universe of tasks, pick the most valuable one.
But there’s a whole class of problems where the hard part isn’t choosing among known options — it’s discovering that the options exist at all. Scientific research, drug discovery, novel material design, open-ended product strategy: nobody hands you a backlog. You have to generate the hypotheses yourself, test them, and decide which dead ends to abandon.
That’s the Exploration and Discovery pattern, and it’s a fitting place to end this series — because it’s where agents stop being faster autocomplete and start behaving like collaborators that can surprise you.
Pattern #21: Exploration and Discovery#
The Problem#
Most agentic patterns we’ve covered are fundamentally exploitative: they take a known objective and execute it efficiently. Routing dispatches to the right specialist. Planning decomposes a goal into steps. RAG retrieves what already exists. All of them assume the space of good answers is more or less known, and the job is to navigate it well.
Discovery is the opposite. The space of good answers is unknown, possibly large, and most of it is worthless. A single LLM call asked to “propose a novel research direction” will happily generate something plausible — and that’s exactly the trap. Plausible is not novel, and it’s not rigorous. One model, talking to itself, has no adversary. It will rationalize its own weak ideas because nothing in the loop is incentivized to tear them apart.
Real discovery has structure for a reason. A research lab doesn’t let one person generate a hypothesis and grade their own work. There’s a division of labor: someone surveys the prior art, someone designs the experiment, several independent reviewers attack it from different angles, and a senior figure weighs the conflicting feedback into a decision. The friction is the point. That adversarial structure is what separates a discovery process from a brainstorm.
The Solution#
The script for this chapter models exactly that lab as a LangGraph workflow. The state carries the research artifacts as they accumulate, with the reviews using the same reducer trick we saw back in Parallelization so that independent reviewers don’t trample each other’s writes:
class DiscoveryState(TypedDict):
research_topic: str
literature_review: str
experimental_plan: str
reviews: Annotated[List[str], operator.add]
final_report: strThe workflow has four distinct roles. First, a PostDoc agent does the groundwork in two sequential steps — survey the literature, then turn it into a concrete plan. Note the temperature: 0 for the review (we want grounded, factual surveying) and 0.3 for the plan (a little creative latitude where it matters):
def literature_review(state: DiscoveryState) -> dict:
"""PostDoc agent: conducts initial literature review."""
llm = get_llm(temperature=0)
prompt = ChatPromptTemplate.from_messages([
("system",
"You are a postdoctoral researcher. Conduct a brief literature review on "
"the given topic. Identify 3-4 key papers/findings and gaps in current research."),
("user", "Literature review for: {research_topic}")
])
chain = prompt | llm | StrOutputParser()
review = chain.invoke({"research_topic": state["research_topic"]})
return {"literature_review": review}Then comes the part that makes this discovery and not just generation: three independent reviewers, each with a different obsession. One is a harsh methodologist hunting for missing controls and statistical holes. One only cares about impact and significance. One only cares about novelty — is this genuinely new, or an incremental retread? They run in parallel and each appends its critique:
def reviewer_novelty(state: DiscoveryState) -> dict:
"""Reviewer 3: focus on novelty and originality."""
llm = get_llm(temperature=0)
prompt = ChatPromptTemplate.from_messages([
("system",
"You are a peer reviewer focused on novelty and originality. "
"Assess whether the approach is truly novel or incremental. "
"Suggest how to differentiate from existing work.\n\n"
"Plan:\n{experimental_plan}"),
("user", "Review this research plan for novelty.")
])
chain = prompt | llm | StrOutputParser()
review = chain.invoke({"experimental_plan": state["experimental_plan"]})
return {"reviews": [f"[Reviewer 3 - Novelty]\n{review}"]}Finally, a Professor agent reads the plan and the full set of conflicting reviews, and synthesizes them into a verdict — accept, revise, or reject — with concrete next steps. This is the role that matters most: it doesn’t average the reviews, it adjudicates them.
The graph wiring makes the lab’s choreography explicit — sequential where there’s a real data dependency, parallel where there isn’t, fan-in where judgment is needed:
# Sequential: literature → plan
builder.add_edge(START, "literature_review")
builder.add_edge("literature_review", "formulate_plan")
# Parallel: all three reviewers
builder.add_edge("formulate_plan", "reviewer_experimental")
builder.add_edge("formulate_plan", "reviewer_impact")
builder.add_edge("formulate_plan", "reviewer_novelty")
# Fan-in: all reviewers → professor
builder.add_edge("reviewer_experimental", "professor_synthesis")
builder.add_edge("reviewer_impact", "professor_synthesis")
builder.add_edge("reviewer_novelty", "professor_synthesis")
builder.add_edge("professor_synthesis", END)This is the whole pattern in one graph: it’s Planning, Parallelization, Multi-Agent, and Reflection composed into a single discovery loop. The adversarial reviewers are the Reflection pattern turned into a panel; the parallel fan-out is Parallelization; the role specialization is Multi-Agent. Discovery isn’t a new primitive — it’s an architecture built from the patterns we’ve spent twenty chapters assembling.
Why This Matters#
This is the pattern behind the systems that actually generate value beyond automation:
- Scientific hypothesis generation: propose research directions, then subject them to adversarial peer review before a human ever reads them. The example topic in the script — using LLMs for automated hypothesis generation — is delightfully recursive.
- Drug and materials discovery: enumerate candidate compounds or structures, score them against multiple independent criteria, and let an adjudicator rank what’s worth synthesizing in a real lab.
- Strategy and product exploration: generate non-obvious market moves and have specialist critics (legal, financial, competitive) shred them before they reach a decision-maker.
- Red-teaming and security research: one agent invents attack hypotheses, others probe for flaws and feasibility.
The honest trade-offs are steep, and you should not romanticize this pattern. It’s expensive. This single run fires six LLM calls in a fixed pipeline — and real discovery is iterative, so you’ll often loop the professor’s “revise” verdict back into a new plan, multiplying the cost. It’s unbounded. Open-ended exploration has no natural stopping point; without budgets, iteration caps, or a novelty threshold, an agent will happily explore forever and bill you for it. And the critics are only as good as their prompts. Three “reviewers” backed by the same base model share the same blind spots — genuine diversity of critique is hard to fake with role-play alone, and that’s the ceiling on how much real novelty this pattern can surface.
Rule of thumb: reach for Exploration and Discovery only when the value is in finding unknown good answers, not executing known ones. If you already know what good looks like, every other pattern in this series is cheaper and more reliable. Discovery is the most powerful tool in the box and the easiest to waste money on.
The Bigger Picture#
This is post #21 — the last — in my series documenting Antonio Gulli’s Agentic Design Patterns. All credit for the conceptual framework goes to him; my contribution has been to turn each chapter into clean, runnable Python you can clone and ship.
What strikes me looking back is that Exploration and Discovery isn’t a standalone trick — it’s a capstone. It only works because it stands on everything else: the Planning that decomposes the work, the Parallelization that runs the critics concurrently, the Multi-Agent role specialization that gives them distinct viewpoints, the Reflection that turns critique into a feedback loop. The whole catalogue composes.
All the code from this post is in my repository: carlosprados/Agentic_Design_Patterns, specifically under 21_Exploration_and_Discovery/. It runs with uv run, supports Gemini and Ollama through the shared get_llm() abstraction, and is ready to fork.
What’s Next#
There is no next chapter — this is where the series ends. Twenty-one patterns, from Prompt Chaining all the way to Exploration and Discovery, each one refactored from concept into production-ready Python.
If you’ve followed along from the start, thank you. The point was never any single pattern in isolation — it was to build the vocabulary to reason about agentic systems as compositions of well-understood pieces, instead of one giant prompt and a prayer. That vocabulary is the real takeaway.
Go build something that surprises you. And if you ship something using these patterns, I’d genuinely like to hear about it.
Stay tuned — the series is done, but the work isn’t.

