Tongyi DeepResearch

repository·main·Indexed 12 days ago

https://github.com/alibaba-nlp/deepresearch

An agentic large language model (30.5B parameters) designed for long-horizon, deep information-seeking tasks. It features AgentFounder-30B, which utilizes Agentic Continual Pre-training (Agentic CPT) to enhance research capabilities, and AgentScaler, a framework for scaling general agentic intelligence through simulated environments. The system supports ReAct and IterResearch-based 'Heavy' modes, with context lengths up to 128K.

Tokens
31.5K
Snippets
53
Records
146
Agent score
98%

What's inside Tongyi DeepResearch

  1. Overview of AgentFounder and Agentic CPT

    main

    AgentFounder is a deep research agent model (specifically AgentFounder-30B) that utilizes Agentic Continual Pre-training (Agentic CPT). Unlike standard post-training, Agentic CPT integrates a systematic agentic training data synthesis method into the training pipeline to enhance research capabilities.

    The core approach involves a two-stage continual pre-training strategy designed to scale agent performance through specialized data synthesis and extended context handling.

  2. Overview of National Teaching Innovation Competitions

    main

    The National Higher Education Teaching Innovation Competition (Innovation Competition) and the National Young Teachers Teaching Competition (Youth Competition) are high-level national teaching contests in China. They serve as engines for higher education reform, focusing on 'Four New' construction (New Engineering, New Medicine, New Agriculture, New Liberal Arts) and 'Curriculum Ideology and Politics' (课程思政).

    Key Characteristics:

    • Organizational Structure: The Innovation Competition is guided by the Ministry of Education's Higher Education Department and organized by the China Association of Higher Education. It follows a three-tier selection process: School $\rightarrow$ Provincial $\rightarrow$ National.
    • Grouping/Tracks:
      • Innovation Competition: Grouped by professional title (Senior, Associate, Intermediate/Below) and includes 7 tracks such as 'Four New' construction, Basic Courses, Curriculum Ideology, and Industry-Education Integration.
      • Youth Competition: Grouped by discipline (Liberal Arts, Science, Engineering, Medicine, and Ideological/Political Special Group).
    • Core Standards: Competitions emphasize the 'Two Qualities and One Degree' (两性一度) principle: High-orderfulness (高阶性), Innovativeness (创新性), and Challenge (挑战度).
  3. Overview of AgentScaler

    main

    AgentScaler is a framework designed to advance general agentic intelligence by scaling up environments. It addresses the challenges of principled environment scaling and effective agent training through a two-phase approach:

    1. Environment Construction and Scaling: Automatically constructs heterogeneous, fully simulated environments to systematically broaden the space of function-calling scenarios.
    2. Agent Experience Learning: A two-phase fine-tuning strategy where agents are first endowed with fundamental agentic capabilities and then specialized for domain-specific contexts.

    AgentScaler has demonstrated significant enhancements in function-calling capabilities on benchmarks such as τ-bench, τ2-Bench, and ACEBench.

  4. Overview of WebAgent models and features

    main

    WebAgent is a collection of models and frameworks developed by Tongyi Lab (Alibaba Group) for information-seeking tasks. The ecosystem includes several specialized agents:

    • WebWatcher: A vision-language deep research agent capable of using tools like Search, Visit, ImageSearch, and CodeInterpreter.
    • WebShaper: Focuses on agentic data synthesis via information-seeking formalization, achieving high scores on GAIA and WebWalkerQA.
    • WebSailor: An agentic search model specialized in complex information seeking through extended thinking and post-training methodologies (e.g., DUPO).
    • WebDancer: A native agentic search reasoning model using the ReAct framework for autonomous information seeking and long-horizon tasks.
    • WebWalker: A benchmark and multi-agent framework for web traversal.
  5. Overview of WebShaper

    main

    WebShaper is a formalization-driven data synthesis method designed for training information-seeking (IS) agents. Unlike traditional methods that retrieve information first and then synthesize data, WebShaper establishes a task formalization first, then collects information and synthesizes QA pairs based on that formalization. This approach aims to prevent redundancy and reasoning shortcuts in synthesized datasets.

    Key components include:

    • Agentic Expander: An iterative component that generates and validates questions aligned with the task formalization.
    • Layer-wise Structure: An expansion paradigm that traverses leaf constants and replaces them with variables to ensure the model must strictly seek information and reason through all variables to reach a target.
  6. Overview of WebSailor-V2

    main

    WebSailor-V2 is a post-training pipeline designed for web agents, covering data construction, Supervised Fine-Tuning (SFT), and Reinforcement Learning (RL).

    Key components include:

    • SailorFog-QA-2: A novel dataset built from a densely interconnected knowledge graph designed to introduce uncertainties and foster sophisticated reasoning.
    • Dual-environment RL framework: A training methodology that combines a high-fidelity simulator (for rapid, low-cost iteration) with a managed real-world environment (for stable final policy training) using a symbiotic data-policy feedback loop.
  7. Overview of WebLeaper Framework

    main

    WebLeaper is a data and training framework designed to improve the effectiveness and efficiency of web agents during information-seeking (IS) tasks. It addresses the issue of agents performing shallow, meandering searches by using entity-intensive tasks and efficiency-aware training.

    Core Components

    • Task Synthesis: Creates three data variants to increase reasoning complexity:
      • Basic: Single-source, high-density tasks from a single table.
      • Union: Multi-source tasks requiring fusion across different sources.
      • Reverse-Union: Tasks that provide attribute-level clues to force deduction before searching, preventing keyword shortcuts.
    • Two-Stage Training:
      1. SFT (Supervised Fine-Tuning): Uses trajectories filtered by Information-Seeking Rate (ISR) and Information-Seeking Efficiency (ISE).
      2. RL (Reinforcement Learning): Uses a Hybrid Reward System and Group Relative Policy Optimization (GRPO) to refine the policy.
    • Key Metrics:
      • ISR (Information-Seeking Rate): The fraction of required entities retrieved.
      • ISE (Information-Seeking Efficiency): The number of target entities discovered per action step.
  8. Explore the Tongyi DeepResearch Agent Family

    main

    Tongyi DeepResearch is part of an extensive family of research agents focused on web traversal, autonomous information seeking, and long-horizon reasoning. Key research papers in this family include:

    • WebWalker: Benchmarking LLMs in Web Traversal
    • WebDancer: Towards Autonomous Information Seeking Agency
    • WebSailor: Navigating Super-human Reasoning for Web Agent
    • WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization
    • WebWatcher: Vision-Language Deep Research Agent
    • WebResearcher: Long-Horizon Agents
    • ReSum: Context Summarization for Long-Horizon Search
    • WebWeaver: Structuring Web-Scale Evidence
    • WebSailor-V2: Synthetic Data and Scalable RL
    • AgentFold: Proactive Context Management
    • WebLeaper: Efficient Info-Rich Seeking
    • BrowseConf: Confidence-Guided Test-Time Scaling
    • ParallelMuse: Agentic Parallel Thinking
    • AgentFrontier: ZPD-Guided Data Synthesis
    • Nested Browser-Use Learning
  9. Leading Research Teams in Quantum Computing

    main

    The following research groups and collaborations are identified as leaders in the field of quantum computing, categorized by their primary technical focus:

    • Hardware-Efficient Error Correction: Caltech-AWS Joint Team. They developed the Ocelot chip, which uses cascaded bosonic (cat) qubits to address quantum error correction challenges.
    • Full-Stack Fault-Tolerant Engineering: Quantinuum Core Team. They have achieved milestones including fault-tolerant non-Clifford gates with logical error rates lower than physical error rates, high-fidelity logical magic state preparation via code switching, and end-to-end error-corrected quantum chemistry calculations.
    • Quantum Software & Compilation: Oxford-Cambridge-Quantinuum Axis. This group focuses on the ZX-calculus, a graphical quantum computing language that serves as the core of Quantinuum's TKET compiler.
    • Theoretical Foundations: Perimeter Institute Theoretical Cluster. Research spans quantum error correction, quantum gravity, high-dimensional quantum systems (qudits), and AI-driven quantum technologies (e.g., applying LLM structures to quantum simulations).
    • Modular Quantum Computing: ETH Zurich-PSI Quantum Hub. Supported by IARPA, they manage the SuperMOOSE (superconducting) and MODULARIS (trapped-ion) projects to develop scalable, modular quantum architectures.
  10. Challenges in Machine Learning for Material Optimization

    main

    Applying machine learning to material element composition optimization faces three primary categories of challenges:

    1. Data Scarcity and Quality Issues

    • Data Homogeneity Bias: Most models rely on high-throughput DFT (Density Functional Theory) databases (e.g., Materials Project, OQMD). Because these are based on the same theoretical approximations (like PBE functionals), models may fail to capture complex experimental dynamics, defects, or environmental effects.
    • Small Sample Sizes: Emerging or complex material systems often have very small datasets (e.g., some perovskite thermal expansion datasets contain only ~137 points), making deep learning prone to overfitting.
    • Data Redundancy and Overestimation: Large databases often contain chemically similar materials. Using random splits for training/testing can lead to artificially high performance (e.g., $R^2 > 0.9$). Performance often drops significantly when evaluated on true out-of-distribution (OOD) sets (e.g., Roost model $R^2$ dropping from 0.92 to 0.53).

    2. Model Interpretability and Trustworthiness

    • Black-Box Nature: Deep learning models lack transparency, which hinders scientific discovery and industrial trust. There is a risk of "shortcut learning," where models exploit correlations in data distribution that lack physical causation.
    • Explainable AI (XAI) Solutions: Methods like SHAP (SHapley Additive exPlanations) and LIME (Local Interpretable Model-agnostic Explanations) are used for feature attribution to quantify how much each input feature contributes to a prediction. Attention mechanisms are also used to visualize model focus.
    • Uncertainty Quantification: Most models lack reliable error estimation and mechanisms to determine the domain of applicability, which is critical for guiding experimental validation.

    3. Computational and Deployment Bottlenecks

    • High-Dimensional Search Spaces: The combinatorial space of material compositions (e.g., a 10-component polymer) is massive. While ML surrogate models are faster than finite-element analysis (e.g., evaluating $>10^5$ candidates in ~4 hours vs. weeks), the computational cost remains significant.
    • Software Dependency and Maintainability: Models often depend on specific Python environments and library versions, leading to usability degradation over time.
    • Lack of Experimental Feedback Loops: Most research is limited to the "prediction" stage. A complete discovery cycle requires integration with autonomous synthesis platforms (e.g., ChemOS, Chemputer) to create a "predict-synthesize-characterize-feedback" loop.
  11. Research Landscape: Machine Learning in Materials Science

    main

    The research in machine learning and deep learning for materials science is categorized into three primary geographical and institutional hubs, each focusing on different methodologies for material discovery and optimization:

    1. US Universities and National Laboratories

    Focuses on high-throughput computing, generative AI, and large-scale databases:

    • Purdue University (SCALE Program): Combines Generative AI (e.g., diffusion models) with Active Learning to optimize microelectronic material compositions (e.g., high-performance alloys).
    • Northwestern University (CHiMaD): Integrates First-principles calculations (DFT) with machine learning for rapid identification of quantum materials.
    • Auburn/Utah (Prasanna Balachandran): Emphasizes Uncertainty Quantification, Bayesian learning, and exploration-exploitation learning for high-dimensional material space exploration.
    • UIUC (OQMD Team): Provides high-throughput DFT repositories (Open Quantum Materials Database) that underpin most supervised ML pipelines.
    • Lawrence Berkeley National Lab (Materials Project): Offers large-scale datasets and online platforms for predicting electronic, thermodynamic, and mechanical properties.
    • MIT & DeepMind: Leading the use of Graph Neural Networks (GNN) (e.g., CGCNN) and large-scale foundation models (e.g., MACE, GNoME) for predicting material stability and inverse design.

    2. European Research Institutions

    Focuses on functional modeling and surface chemistry:

    • Queen Mary University of London: Uses Artificial Neural Networks (ANN) to predict functional properties of ceramics (e.g., dielectric constants) directly from composition.
    • University of Warwick (Reinhard Maurer): Integrates ML with first-principles methods to study heterogeneous catalysis and surface chemistry.
    • University College London (Gaultois): Pioneered data-driven thermoelectric databases and web-based ML tools.
    • University of Oxford (Antunes & Butler): Utilizes Attention-based deep learning models for direct composition-to-property mapping in thermoelectric materials.

    3. Asian Research Teams

    Focuses on industrial applications and closed-loop workflows:

    • Harbin Institute of Technology: Developed a closed-loop workflow combining sparse experimental data, ML models, Maxwell-Garnett theory, and electromagnetic simulations for microwave absorbing materials.
    • Shanghai University (Wencong Lu): Focuses on materials informatics, data mining, and performance optimization.
    • Daicel Allnex (S. Muroga): Developed a Multimodal Deep Learning (MDL) framework combining GAN-based generative models (optical microscopy, IR spectra, Raman spectra) with regression networks for polymer composites.
    • University of Science and Technology of China (Shen Baolong): Applies Active Learning to reduce DFT calculation costs in thermoelectric material discovery.
    • Tsinghua University (Zhang Yingying): Focuses on ML-assisted design for graphene-based flexible materials.
  12. What is the Reverse-Ephemeris Lunar Navigation System?

    main

    The Reverse-Ephemeris Lunar Navigation System is a low-cost navigation concept where the roles of the satellite and the receiver are reversed.

    • Mechanism: Surface-based S-Band transceivers (2,400 – 2,450 MHz, 10 W) transmit signals to a small constellation of three satellites in frozen elliptical lunar orbits.
    • Operation: The satellites act as fixed reference points because their orbits (ephemerides) are precisely known. Surface assets determine their position by measuring the signal travel time from the surface to the satellite.
    • Capacity: Designed to provide continuous coverage for up to 300 simultaneous users over 1.8 MHz of bandwidth.
    • Advantage: Highly cost-effective and robust against jamming compared to traditional satellite-to-ground systems.