Datasets
A catalogue of 15 datasets for training and evaluating models on the AETRON network: licence, size, format and the task types each one pairs with.
- FineWeb · Hugging Face
15T-token English web corpus from 96 CommonCrawl snapshots — the open pretraining standard.
- OpenAssistant Conversations 2 (OASST2) · OpenAssistant / LAION
Human-written assistant conversations in 35 languages with quality rankings, incl. Russian.
- Tulu 3 SFT Mixture · Allen Institute for AI (Ai2)
Ai2's 939K-sample SFT mixture behind Tulu 3 — an open recipe for state-of-the-art post-training.
- UltraFeedback · OpenBMB / Tsinghua
64K prompts with 256K GPT-4-annotated completions — the standard preference set for DPO.
- HH-RLHF · Anthropic
Anthropic's helpfulness & harmlessness preference pairs plus red-teaming dialogues.
- MMLU · CAIS / Dan Hendrycks
57-subject multiple-choice benchmark — the de-facto exam for LLM knowledge.
- GSM8K · OpenAI
8.5K grade-school math word problems with step-by-step solutions — core reasoning benchmark.
- MS MARCO · Microsoft
Million-scale Bing question-answering and passage-ranking corpus for retrieval and RAG.
- Common Voice 17.0 · Mozilla Foundation
Mozilla's crowdsourced speech corpus — 31K hours, 120+ languages, CC0.
- LibriSpeech ASR · OpenSLR
1,000 hours of read English audiobooks — the classic ASR benchmark.
- LJ Speech 1.1 · Keith Ito / LibriVox
13.1K clips of a single English speaker — the standard TTS corpus.
- COCO 2017 · COCO Consortium
123K images with captions, bounding boxes and segmentation masks.
- ImageNet-1K (ILSVRC 2012) · ImageNet (Stanford / Princeton)
The canonical 1.28M-image, 1,000-class classification benchmark. Research-only.
- DiffusionDB · Georgia Tech / Polo Club
14M Stable Diffusion images with the exact prompts that produced them.
- LLaVA-Instruct-150K · LLaVA Team (Liu et al.)
158K GPT-4-generated visual instructions over COCO images. Research-only.