Overview of C-MTEB Tasks and Datasets
masterC-MTEB (Chinese Massive Text Embedding Benchmark) provides a comprehensive suite of evaluation tasks for Chinese embedding and reranking models. The tasks are categorized into several types:
- Retrieval: Includes datasets like
T2Retrieval,MMarcoRetrieval,DuRetrieval,CovidRetrieval,CmedqaRetrieval,EcomRetrieval,MedicalRetrieval, andVideoRetrieval. For these tasks, 100,000 candidates (including ground truths) are sampled from the corpus to manage inference costs. - Reranking: Includes
T2Reranking,MMarcoReranking,CMedQAv1, andCMedQAv2. - PairClassification: Includes
OcnliandCmnli. - Clustering: Includes
CLSClusteringS2S,CLSClusteringP2P,ThuNewsClusteringS2S, andThuNewsClusteringP2P. - STS (Semantic Textual Similarity): Includes
ATEC,BQ,LCQMC,PAWSX,STSB,AFQMC, andQBQTC. - Classification: Includes
TNews,IFlyTek,Waimai,OnlineShopping,MultilingualSentiment, andJDReview.