The following datasets provide large-scale, structured, or multilingual knowledge for building and enriching knowledge graphs:
- BabelNet: A multilingual encyclopedic dictionary and semantic network with ~16 million Babel synsets.
- Wikidata: A collaborative, multilingual database providing structured data for the Wikimedia movement.
- Google Knowledge Graph: Millions of entries describing real-world entities (people, places, things).
- Freebase: A large-scale knowledge base (acquired by Google and used in Google Knowledge Graph).
- DBpedia: Structured content extracted from Wikimedia projects.
- XLore: A large-scale English-Chinese bilingual knowledge graph.
- The GDELT Project: Monitors global news in 100+ languages to identify entities, events, and themes.
- YAGO: A semantic knowledge base derived from Wikipedia, WordNet, and GeoNames (~10M entities, 120M facts).
- Zhishi.me: Knowledge Graph data from major Chinese encyclopedias (Baidu Baike, Hudong Baike, Chinese Wikipedia).
- NELL (Never-Ending Language Learner): Continuously extracts facts from web pages.
- Golden Protocol: A decentralized, open, and transparent Web 3 knowledge graph.