Assess Deep Research production-readiness
mainThe Deep Research feature has been evaluated through two rounds of testing (R1 and R2) to measure the impact of grounding refinements. The current verdict is READY-WITH-CAVEATS.
While the refinement successfully eliminated the primary failure mode of fabricated sources (improving from 0.0 to ~27.5 unique sources per report), users should be aware of the following residual risks:
- Marginal Fabrication: A 35B local model may still produce plausible-looking but unverifiable specifics, such as future-dated CVE IDs or arXiv IDs. Exact identifiers, versions, and figures should always be verified.
- Latency: If subagents hang, the watchdog mechanism will kill them, which can increase total execution time (up to ~54 minutes in serial mode, though parallel mode in production mitigates this).
- Report Thinness: If research subagents are killed by the watchdog, the resulting report may be thinner than expected due to missing information slices.
- Judge Unreliability: Do not rely on local 35B models to automatically gate quality via a 'judge' metric, as they can be unreliable and may penalize honest reports that flag information gaps.