A Deep Research Workflow for Complex Questions
An end-to-end workflow for multi-source research that needs transparent methods, synthesis, uncertainty, and review.
An end-to-end workflow for multi-source research that needs transparent methods, synthesis, uncertainty, and review.

A deep research workflow is appropriate when a question spans multiple sources, subtopics, definitions, or competing explanations. It requires a written scope, decomposition, transparent search plan, source register, claim ledger, synthesis across evidence, and an audit of the final output. More searching is not automatically deeper research; better traceability and reasoning are.
The answer will influence a consequential decision.
No single authoritative source resolves it.
The question contains multiple time periods, jurisdictions, or populations.
Sources use conflicting metrics or definitions.
The output must explain tradeoffs, causes, or uncertainty.
A reusable report, dataset, or evidence packet is required.
Write the decision, audience, and as-of date.
Define key terms and exclusions.
Break the question into factual, comparative, causal, and evaluative parts.
Set source hierarchy and stopping rules.
Choose the final output before gathering everything.
Use separate lanes for official sources, original research, data, implementation evidence, criticism, and current updates. This produces genuine source diversity instead of ten pages repeating one announcement. For scholarly review reporting, PRISMA offers specific guidance and flow tools; use it when the review type fits rather than as a decorative label.
Atomic claim and subquestion
Supporting evidence and locator
Source type and authority
Scope, date, definitions, and units
Contradicting evidence
Inference required
Confidence and unresolved gap
Organize by question, not source. Explain where evidence converges, then test whether apparent conflict comes from dates, samples, denominators, methods, or incentives. Keep causal language proportional to study design. The National Academies stresses scrutiny, repeated testing, and explicit uncertainty as foundations of scientific knowledge.
Lead with the decision-relevant answer and confidence.
Describe scope and method briefly.
Tie findings to claim-level sources.
Show material contradictions and missing data.
Use charts only for verified numeric patterns.
Separate recommendation criteria from factual findings.
Run a final citation, calculation, freshness, and reasoning review.
Rixx can support this loop with cited web research, supported documents, follow-ups, charts, reports, and organized Insights where available. Deep research still depends on the researcher’s source choices and review.
Suppose the question is whether a city should encourage a particular transport policy. The research cannot begin with a global list of benefits. It needs the city’s objective, baseline, affected population, implementation constraints, and decision horizon. Subquestions may cover safety, travel time, access, cost, emissions, enforcement, and distributional effects. Each requires different evidence and may use a different unit.
The source plan could include local administrative data, enacted rules, transport-agency methodology, comparable-city evaluations, peer-reviewed studies, and affected-community evidence. A finding from another city is not automatically transferable. The synthesis should state which conditions match, which do not, and whether an observed association supports a causal claim. The recommendation can then offer conditions or a pilot instead of a false yes-or-no certainty.
Scope gate: stakeholders agree on the question, definitions, and exclusions.
Retrieval gate: each subquestion has suitable source coverage or a recorded gap.
Evidence gate: consequential claims have locators, dates, and scope notes.
Synthesis gate: disagreement is explained rather than averaged away.
Analysis gate: calculations are reproducible and units consistent.
Output gate: recommendations follow from stated criteria.
Review gate: an appropriate person has checked high-risk conclusions.
Running one broad query and treating the top results as the source universe.
Collecting many sources without recording what each contributes.
Summarizing documents before reading methods and limitations.
Confusing source diversity with multiple pages repeating one dataset.
Choosing a conclusion before defining decision criteria.
Using precise scores to hide qualitative judgments.
Generating a long report that has no claim-level evidence map.
Stop when new credible sources mainly repeat established findings, the important contradictions have been investigated, and remaining gaps are explicit enough for the decision maker. Also stop when evidence cannot answer the question. More searching cannot repair an unavailable counterfactual, an undefined metric, or missing local data. The honest output may be a research plan rather than a recommendation.
Keep a lightweight log of search date, query, source or database, filters, and major inclusion decisions. For web research, note pages that are dynamic or likely to change. For data, save the source release, extraction date, and transformations. The goal is not laboratory-level reproduction for every question; it is enough transparency for another researcher to understand how the evidence set was formed and where it may be incomplete.
Research owner: defines scope and accepts the final method.
Source reviewer: checks authority, identity, currency, and coverage.
Domain reviewer: challenges interpretation and missing context.
Data reviewer: verifies calculations, units, and charts.
Editor: tests whether wording matches evidence and audience.
Decision owner: accepts residual uncertainty for the intended action.
One person may hold several roles, but naming them exposes missing checks. A solo researcher can perform separate passes rather than pretending simultaneous drafting and verification are independent. For sensitive or consequential work, independence matters more: a specialist should review the claims that require their expertise.
Deep research also needs a clear archive boundary. Retain what the decision requires, respect licenses and privacy, and avoid collecting sensitive material without purpose. Record inaccessible evidence rather than pretending it was reviewed. If a later researcher can see what was searched, selected, rejected, calculated, and left unresolved, the work remains useful even when its conclusion changes.
Depth is not the number of pages produced. It is how well the conclusion survives inspection from question to source.
A short transparent conclusion with an evidence register is deeper than a long report whose claims cannot be reconstructed. Preserve that distinction at every handoff.
Before using this guidance, return to the actual decision and test it against deep research workflow, deep research process, multi-source research, and AI deep research. Record which evidence is direct, which conclusion is inferred, which facts can change, and who will review the result. Check the strongest counterexample, preserve source dates and definitions, and stop when missing evidence could reverse the decision. A useful output should remain understandable without hidden chat context and correctable when a source changes. Do not convert an unavailable fact into an estimate, an example into a testimonial, or a product direction into a promise.