The five patterns
Fan-out and merge
The single biggest quality difference between a thin answer and a good one. One query is one search; a real question deserves 5–30.1
Decompose before retrieving
Split the question into subquestions — mechanism, population, outcome, competing explanations. Generate several phrasings per subquestion covering synonyms and adjacent terminology.
2
Pace them to your rate limit
Every query is independent, but there is a one-query-per-second rate limit on every plan except Enterprise. Unbounded parallelism gets you 429s, not speed — run searches sequentially and build the pacing into your system. See rate limits below.
3
Deduplicate on DOI
Merge on
doi, keeping the first-seen record plus the list of queries that found it. How many queries surfaced a paper is a useful relevance signal in itself.4
Keep the pre- and post-merge counts
You will need both to describe your coverage, and a sharp drop tells you your phrasings were near-duplicates.
Filter ladder
Exhaustive sweep
Pagination and page sizes above the default of 20 require a paid plan or an Enterprise API key.
page_size caps at 300 on Pro and Teams, 750 on Deep, and higher on Enterprise keys. A request over your ceiling is clamped rather than rejected — check the page_size echoed back in the response.Rate limits and result caps
Two separate ceilings shape every workflow here, and confusing them is the most common cause of a pipeline that quietly under-reports. Rate limit — how fast you may ask. One search per second on every plan except Enterprise:
Retrieving details for a paper you have already found is not the bottleneck — searching is.
Result cap — how much comes back. A response that says it found more papers than it returned has been capped by your plan. That distinction matters: you can only cite what was shown, and a thin result set may reflect your plan ceiling rather than a genuine gap in the literature. Parse both numbers, and when they differ, say so in your output rather than reporting the shown count as the whole literature.
Handling 429s. The two kinds mean different things:
- Per-second rate limit — there is a one-question-per-second limit on every plan except Enterprise. Account for it in whatever you are building rather than firing many searches in parallel: pace your calls, honour
retry-after, and add jitter. Retrying helps. - Monthly search allowance — you have used the searches included in your plan. Retrying does not help. Turn on additional usage in your account settings to keep searching past the included allowance.
Grounding and auditability
Every workflow in the gallery ends in something a person will act on — a citation in a manuscript, a target decision, a claim that ships. These five rules are what keep that output trustworthy, and they are the ones the Consensus-built skills enforce most strictly.1
Cite only what came back in this session
Never supplement results with model knowledge, a previous conversation, or a paper you happen to know. A fabricated citation in an academic or regulatory context costs far more than a thin result set. If you include something not retrieved in this session for context, label it explicitly and exclude it from every count.
2
Count three things separately
Queries sent, results returned, and results cited. These are different numbers and conflating them is how a five-paper review gets described as comprehensive. Track all three from the first search, not retroactively.
3
Detect the plan cap and report it
When a response reports finding more papers than it returned, your plan capped it. Only the returned papers are citable. Surface it: a sparse section may reflect your ceiling rather than a genuine gap in the literature, and the reader cannot tell the difference unless you say which it was.
4
Confirm each result before moving on
A step is not done because you sent the request — it is done when the response is back and contains data. On failure, wait 3 seconds and retry once; after 3 consecutive failures across any combination of calls, stop and report. Never silently skip a failed search, and never present partial coverage as complete.
5
Require a retrievable link for every citation
No URL or DOI from this session means not citable. Enforce this at write time rather than at review time — it is the single check that catches the most problems.
The audit log
End any workflow that produces a deliverable with a table of what it actually did:
Follow it with the totals, any plan cap you detected, and any search that failed after its retry. This is what lets a reader — or a reviewer, or a mentor — judge how much weight the output deserves, and it is cheap to produce if you have been counting from the start.
What makes a good Consensus use case
The question has an evidence base
The question has an evidence base
Consensus searches peer-reviewed literature. It answers “what does the published evidence say about X” extremely well. It is not a source for your internal experiment data, patent full text, regulatory filings, or private clinical data — join those to Consensus results rather than expecting Consensus to hold them.
The population can be narrowed with filters
The population can be narrowed with filters
The filters are where accuracy comes from. A question you can express as “human RCTs, 100+ participants, Q1 journals, since 2020” comes back far cleaner than an unfiltered keyword search. If you cannot describe the study population you want, spend a turn deciding it before you search.
The output is verifiable
The output is verifiable
Every result carries
title, authors, journal_name, publish_year, citation_count, doi, and a url back to the paper on Consensus. Good use cases end in something a human can audit line by line — an evidence table, a cited memo, a screening decision log. Ones that end in an uncited assertion waste the citation trail.It decomposes into many searches
It decomposes into many searches
Nearly every high-value workflow in the gallery is a fan-out. If your design issues one search and stops, you are leaving most of the quality on the table. Breadth comes from the number of queries, not from how fast you fire them — pace them to your tier’s limit.
Poor fit: single-fact lookups and non-research questions
Poor fit: single-fact lookups and non-research questions
If the answer is one number from a database, a definition, or today’s news, Consensus is the wrong tool. Reach for it when the answer is contested, cumulative, or needs a citation.
Surface differences worth knowing
Back to the gallery
Browse all 12 use cases by persona.
Search endpoint reference
Every filter, the full response schema, and an interactive playground.