DeepKE Knowledge Extraction Toolkit

repository·main·Indexed 26 days ago

https://github.com/zjunlp/deepke

A deep learning-based knowledge extraction toolkit for knowledge graph construction. DeepKE supports entity, relation, and attribute extraction across various scenarios, including cnSchema, low-resource (few-shot), document-level, and multimodal settings. It provides implementations for multiple model architectures such as CNN, RNN, Capsule, GCN, Transformer, and BERT, and includes examples for Event Extraction (EE) and LLM-based extraction methods like CodeKGC, CPM-Bee, and InstructKGC.

Tokens
76.2K
Snippets
228
Records
429
Agent score
85%

What's inside DeepKE

  1. Overview of the Knowledge Graph Construction Task

    main

    The kg2instruction component aims to build a knowledge graph by extracting specific entity and relation types from a given text based on user-provided instructions.

    In a typical Knowledge Graph Construction task, the system receives:

    1. input: The source text.
    2. instruction: A natural language prompt containing the desired entity types or relation types.

    The system's goal is to output all relation triplets found in the input, formatted according to the specification provided in the instruction (e.g., (head_entity, relation, tail_entity)).

  2. Overview of DeepKE-cnSchema

    main

    DeepKE-cnSchema is an out-of-the-box version of the DeepKE open-source knowledge graph extraction and construction tool. It is specifically designed for the Chinese language and supports the cnSchema standard.

    Key features include:

    • Task Support: Named Entity Recognition (NER), Relation Extraction (RE), and Attribute Extraction.
    • Capabilities: Supports low-resource, long-form, and multi-modal knowledge extraction.
    • cnSchema Support: Capable of extracting 50 relation types and 28 entity types (including common types like Person, Location, City, and Organization, and relations like Ancestry, Birthplace, Nationality, and Dynasty).
    • Framework: Built on PyTorch.
  3. Overview of DeepKE-LLM Models

    main

    DeepKE-LLM supports several large language model families and specialized information extraction models:

    • OneKE: A bilingual (Chinese and English) schema-based information extraction model.
    • LLaMA-series: Supports LoRA fine-tuning and specialized usage like ZhiXi.
    • ChatGLM: Supports LoRA fine-tuning and P-Tuning.
    • MOSS: Supports OpenDelta fine-tuning.
    • Baichuan: Supports OpenDelta fine-tuning.
    • CPM-Bee: Supports OpenDelta fine-tuning.
    • GPT-series: Supports various tasks including Information Extraction, Data Augmentation, and Few-shot Relation Extraction.
  4. Overview of DeepKE

    main
    DeepKE is an open-source knowledge graph extraction and construction tool. Built on PyTorch, it supports low-resource, long-context, and multi-modal knowledge extraction, including tasks such as Named Entity Recognition (NER), Relation Extraction (RE), and Attribute Extraction. The DeepKE-cnSchema version provides out-of-the-box support for entity and relation extraction using the cnSchema standard.
  5. Overview of DeepKE functions

    main

    DeepKE is a knowledge extraction toolkit built on PyTorch designed for low-resource and document-level scenarios. It provides a modular and extensible framework for three primary information extraction tasks:

    1. Named Entity Recognition (NER): Identifying and classifying entities in text.
    2. Relation Extraction (RE): Identifying relationships between entities.
    3. Attribute Extraction (AE): Extracting specific attributes associated with entities.

    The toolkit allows for customizing datasets and models to extract information from unstructured texts.

  6. Understand InstructIE and IEPile datasets

    main

    The InstructionKGC project utilizes two primary bilingual (Chinese and English) instruction datasets for Information Extraction (IE):

    1. InstructIE: A dataset with 30w+ samples covering 12 topic categories. Each record contains:

      • id: Unique identifier.
      • cate: Topic category.
      • text: Input text.
      • relation: List of extracted triplets (head, head_type, relation, tail, tail_type).
    2. IEPile: A large-scale dataset with 200w+ samples (0.32B tokens). Each record contains:

      • task: One of NER, RE, EE, EET, or EEA.
      • source: Origin dataset.
      • instruction: A JSON string containing instruction (task description), schema (list of entity/relation/event types), and input (text).
      • output: A JSON string where keys are schema types and values are extracted contents.
  7. Use deepke.attribution_extraction.standard.tools modules

    main

    The deepke.attribution_extraction.standard.tools package provides a suite of utility modules for standard attribution extraction tasks. These modules cover the full pipeline including data handling, preprocessing, training, and evaluation:

    • dataset: Utilities for managing and loading attribution datasets.
    • metrics: Implementation of evaluation metrics specific to attribution extraction.
    • preprocess: Tools for cleaning and preparing raw text/data for the model.
    • serializer: Functions for saving and loading model states or processed data.
    • trainer: Logic for executing the training loop and managing model optimization.
    • vocab: Utilities for vocabulary management and tokenization mapping.
  8. Understand the Knowledge Graph Construction Task Object

    main

    The task objective is to extract specified entities and relationships from text based on user instructions to construct a knowledge graph.

    An example task involves providing an instruction (defining the role, candidate relation list, and output format) and an input (the text). The system outputs relationship triples in the specified format, such as (head entity, relation, tail entity), or NAN if a relation does not exist.

    instruction="You are an expert specifically trained in extracting relation triples. Given the candidate relation list: ['achievement', 'alternative name', 'area', 'creation time', 'creator', 'event', 'height', 'length', 'located in', 'named after', 'width'], please extract the possible head and tail entities from the input below based on the relation list, and provide the corresponding relation triple. If a certain relation does not exist, output NAN. Please answer in the format of (Subject,Relation,Object)\n."
    input="Wewak Airport, also known as Boram Airport, is an airport located in Wewak, Papua New Guinea. IATA code: WWK; ICAO code: AYWK."
    output="(Wewak Airport,located in,Wewak)\n(Wewak,located in,Papua New Guinea)\n(Wewak Airport,alternative name,Boram Airport)\nNAN\nNAN\nNAN\nNAN\nNAN\nNAN\nNAN\nNAN\nNAN"
  9. Understand the InstructKGC Task Object format

    main

    InstructKGC treats Knowledge Graph Construction (KGC) as an autoregressive generation task. The model processes an instruction (a JSON-formatted string) to understand the task intent, schema, and input text, then outputs extracted information in a specified JSON format.

    The instruction string contains three key fields:

    1. 'instruction': A task description (e.g., NER, RE, EE, EET, or EEA).
    2. 'schema': A list of entity types, relation types, or event types to be extracted.
    3. 'input': The source text for extraction.

    The final output JSON object's keys follow the same order as the schema provided in the instruction.

    {
        "task": "NER", 
        "source": "CoNLL2003", 
        "instruction": "{\"instruction\": \"You are an expert in named entity recognition. Please extract entities that match the schema definition from the input. Return an empty list if the entity type does not exist. Please respond in the format of a JSON string.\", \"schema\": [\"person\", \"organization\", \"else\", \"location\"], \"input\": \"284 Robert Allenby ( Australia ) 69 71 71 73 , Miguel Angel Martin ( Spain ) 75 70 71 68 ( Allenby won at first play-off hole )\"}", 
        "output": "{\"person\": [\"Robert Allenby\", \"Allenby\", \"Miguel Angel Martin\"], \"organization\": [], \"else\": [], \"location\": [\"Australia\", \"Spain\"]}"
    }
  10. Understand DeepKE-cnSchema capabilities and performance

    main

    DeepKE-cnSchema is an off-the-shelf version of DeepKE designed for Chinese knowledge graph construction using the CnSchema framework. It supports both Named Entity Recognition (NER) and Relation Extraction (RE) for Chinese text.

    NER Performance

    DeepKE-cnSchema (NER) uses chinese-bert-wwm and chinese-roberta-wwm-ext models.

    • RoBERTa-wwm-ext: F1 score of 0.8310
    • BERT-wwm: F1 score of 0.8197

    RE Performance

    DeepKE-cnSchema (RE) uses chinese-bert-wwm and chinese-roberta-wwm-ext models.

    • RoBERTa-wwm-ext: F1 score of 0.7327
    • BERT-wwm: F1 score of 0.7473
  11. Understand the InstructionKGC task format

    main

    InstructionKGC treats Knowledge Graph Construction (KGC) as an instruction-following autoregressive generation task. The model receives an instruction field containing a JSON-like dictionary string, which consists of three components:

    1. instruction: A natural language task description specifying the model's role and the task to perform.
    2. schema: A dynamic list of labels or event types to be extracted, representing user requirements.
    3. input: The source text for information extraction.

    The model outputs a JSON string where the keys correspond to the schema provided in the input, and the order of keys in the output matches the order in the input schema.

  12. Understand the Instruction-based Knowledge Graph Construction task

    main

    The task involves extracting relevant entities and relations from an input text based on user-provided instructions to construct a knowledge graph. This includes knowledge graph completion, where the model must extract entity-relation triples (ent1, rel, ent2) while identifying missing information.

    Example Task Structure:

    • Instruction: A natural language prompt specifying desired entity types (e.g., {'专业','时间',...}) and relationship types (e.g., {'体育运动','包含行政领土',...}).
    • Input: The source text.
    • Output: A list of triples in the format (head_entity, relation, tail_entity).