Paper2Any

repository·main·Indexed 25 days ago

https://github.com/opendcai/paper2any

A multimodal workflow engine and toolkit designed to transform scientific papers (PDFs, images, text) into editable assets such as diagrams, slide decks, posters, and video scripts. It includes specialized features like Paper2Figure, Paper2PPT, Paper2Rebuttal, and an Image Model Playground, supported by the dataflow-agent LLM component for workflow orchestration.

Tokens
37K
Snippets
74
Records
183
Agent score
82%

What's inside paper2any

  1. Overview of Paper2Any multimodal workflows

    main

    Paper2Any is a toolkit designed for multimodal scientific workflows. It enables users to convert paper PDFs, screenshots, or text into various editable formats, including model diagrams, technical roadmaps, experimental plots, slide decks, academic posters, and video scripts.

    Key capabilities include:

    • Paper2Figure: Generates editable scientific figures (PPTX/SVG).
    • Paper2Diagram / Image2Drawio: Creates editable Draw.io diagrams.
    • Paper2PPT: Converts papers/text/topics into editable slide decks.
    • Paper2Rebuttal: Drafts structured rebuttal responses.
    • PDF2PPT / Image2PPT: Converts PDFs or images into structured slides.
    • Paper2Poster: Generates academic poster layouts from PDFs.
    • Paper2Citation: Explores citations and downstream works via DOI or URLs.
    • Image Model Playground: Provides managed access to image generation models with batching and prompt templates.
  2. Overview of Paper2Any core features

    main

    Paper2Any is a multimodal workflow tool designed to transform scientific papers (PDFs, screenshots, or text) into various editable formats. Key capabilities include:

    • Paper2Figure: Generates editable scientific diagrams (model architectures, technical roadmaps, and experimental data plots) as PPTX.
    • Paper2Diagram / Image2Drawio: Converts papers, text, or images into Drawio diagrams with support for drawio, png, and svg exports.
    • Paper2PPT: Generates editable presentations from papers, text, or themes, supporting long documents and data extraction.
    • Paper2Rebuttal: Drafts structured rebuttal responses and modification suggestions for peer reviews.
    • PDF2PPT & Image2PPT: Converts PDFs with layout preservation or images/screenshots into structured PPTX slides.
    • Paper2Video: Generates video scripts and narration materials for paper explanations.
    • Paper2Poster: Automatically converts paper PDFs into academic posters with customizable layouts and logo injection.
    • Paper2Citation: Tracks citations by author name, DOI, or paper links, including authors, institutions, and representative papers.
    • Knowledge Base (KB): Supports file vectorization, semantic retrieval, and KB-driven generation of PPTs, podcasts, or mind maps.
    • Image Generation Experience: Access to hosted models (e.g., Nano Banana, Image 2) with support for prompt templates, text language control, and batch generation.
  3. Overview of Paper2Video features

    main

    Paper2Video is a tool designed to assist researchers in converting paper content into video scripts or presentation videos.

    Core features include:

    • Video Script Generation: Generates a script suitable for video narration based on the structure of the paper.
    • Keyframe Planning: Suggests visual frames or keyframes for different stages of the video.
    • Multimodal Synthesis (In Development): Aims to automatically synthesize presentation videos by combining Text-to-Speech (TTS) and image generation.
  4. Overview of Paper2Any Workflows

    main

    Paper2Any provides several specialized workflows for transforming academic papers into different media formats. Use the following guide to identify the workflow that matches your input and desired output:

    • Paper2Figure: Converts paper content into academic diagrams, structural charts, flowcharts, and illustrated pages.
      • Inputs: Paper PDFs, abstracts, paragraphs, or technical route descriptions.
      • Outputs: Diagrams, charts, and PPT page assets.
    • Paper2PPT: Converts papers or topics into structured presentations.
      • Inputs: Paper PDFs, topic text, or long-form documents.
      • Outputs: Editable PPT pages and export files.
    • Paper2Video: Converts paper content into a narrated video pipeline, including scripts, voiceovers, and visual pages.
      • Inputs: Paper PDFs, topics, or PPT pages.
      • Outputs: Scripts, subtitles, video clips, and finished videos.
    • Paper2Technical: Extracts methodological details to generate technical route descriptions and structured technical reports.
      • Inputs: Paper PDFs, methodology sections, or experimental descriptions.
      • Outputs: Technical interpretations, process descriptions, and structured reports.
  5. Overview of Paper2PPT features

    main

    Paper2PPT is a tool designed to quickly convert academic papers into structured presentation slides (PPT). Key capabilities include:

    • Intelligent Outline Generation: Automatically extracts core arguments from a paper to generate a PPT outline.
    • Long Document Support: Handles ultra-long papers with automatic pagination and summarization.
    • Figure and Table Extraction: Automatically identifies and extracts images and tables from the paper and inserts them into the corresponding PPT slides.
    • Multiple Templates: Supports Beamer-style templates and various custom PPT templates.
  6. Convert images to editable DrawIO diagrams with Drawio

    main
    The Drawio feature allows you to upload a paper figure or screenshot and convert it into an editable DrawIO canvas. You can then generate model or system diagrams directly within the DrawIO workbench and refine the architecture using chat-based editing and export-ready layouts.
  7. Understand the PaperBanana Methodology Diagram structure

    main

    The PaperBanana framework is visualized as a horizontal flowchart divided into three main functional regions:

    1. Inputs: Consists of Source Context ($S$) and Communicative Intent ($C$).
    2. Linear Planning Phase (Light Blue region):
      • Retriever Agent: Uses Reference Set ($\mathcal{R}$) and Inputs to produce Relevant Examples ($\mathcal{E}$).
      • Planner Agent: Uses Inputs and Relevant Examples to produce an Initial Description ($P$).
      • Stylist Agent: Uses Initial Description ($P$) and Aesthetic Guidelines ($\mathcal{G}$) from the Reference Set to produce an Optimized Description ($P^*$).
    3. Iterative Refinement Loop (Light Orange region):
      • Visualizer Agent: Takes Optimized Description ($P^*$) and Refined Description ($P_{t+1}$) to generate an Image ($I_t$).
      • Critic Agent: Evaluates the Image ($I_t$) against the original Inputs ($S, C$) via Factual Verification to produce a Refined Description ($P_{t+1}$).

    Final Output: The polished Final Illustration ($I_T$).

  8. Generate presentations with Paper2PPT

    main

    Paper2PPT converts papers, text, or topics into polished slide decks. Key capabilities include:

    • AI-Assisted Outline Refinement: Use targeted rewrite prompts to refine outlines at the section and bullet level.
    • Canvas Editing: Edit slide text directly on the canvas while maintaining the deck's theme.
    • Advanced Features: Support for long documents (40+ slides), intelligent table extraction and insertion, and version history for iterative management.
    • Gallery Review: Review the entire multi-page deck before exporting.
  9. Explore Paper2Any Core Features and Capabilities

    main

    Paper2Any provides a suite of tools designed to transform scientific papers into various formats. Key features include:

    • Paper2Figure: Generates editable scientific figures (model architectures, technical roadmaps, experimental data plots).
    • Paper2Diagram (Drawio): Generates Drawio diagrams from text/papers, converts images to Drawio, and supports conversational editing and exports (PNG/SVG).
    • Paper2PPT: Converts papers to editable presentations, supporting Beamer styles, long-form text to PPT, template-based generation, and knowledge-base (KB) driven generation.
    • PDF2PPT & Image2PPT: Converts PDFs or images into editable PPTX files while preserving layouts and performing intelligent image cropping.
    • PPTPolish: Provides intelligent beautification and style transfer for presentations.
    • Knowledge Base (KB) Workflow: Supports file ingestion, vectorization, semantic retrieval, and generating PPTs, podcasts, or mind maps from a knowledge base.
    • Paper2Video: Generates video scripts from papers (scripting and voiceover in progress).