Datasets

A catalogue of 15 datasets for training and evaluating models on the AETRON network: licence, size, format and the task types each one pairs with.

  • FineWeb · Hugging Face

    15T-token English web corpus from 96 CommonCrawl snapshots — the open pretraining standard.

  • OpenAssistant Conversations 2 (OASST2) · OpenAssistant / LAION

    Human-written assistant conversations in 35 languages with quality rankings, incl. Russian.

  • Tulu 3 SFT Mixture · Allen Institute for AI (Ai2)

    Ai2's 939K-sample SFT mixture behind Tulu 3 — an open recipe for state-of-the-art post-training.

  • UltraFeedback · OpenBMB / Tsinghua

    64K prompts with 256K GPT-4-annotated completions — the standard preference set for DPO.

  • HH-RLHF · Anthropic

    Anthropic's helpfulness & harmlessness preference pairs plus red-teaming dialogues.

  • MMLU · CAIS / Dan Hendrycks

    57-subject multiple-choice benchmark — the de-facto exam for LLM knowledge.

  • GSM8K · OpenAI

    8.5K grade-school math word problems with step-by-step solutions — core reasoning benchmark.

  • MS MARCO · Microsoft

    Million-scale Bing question-answering and passage-ranking corpus for retrieval and RAG.

  • Common Voice 17.0 · Mozilla Foundation

    Mozilla's crowdsourced speech corpus — 31K hours, 120+ languages, CC0.

  • LibriSpeech ASR · OpenSLR

    1,000 hours of read English audiobooks — the classic ASR benchmark.

  • LJ Speech 1.1 · Keith Ito / LibriVox

    13.1K clips of a single English speaker — the standard TTS corpus.

  • COCO 2017 · COCO Consortium

    123K images with captions, bounding boxes and segmentation masks.

  • ImageNet-1K (ILSVRC 2012) · ImageNet (Stanford / Princeton)

    The canonical 1.28M-image, 1,000-class classification benchmark. Research-only.

  • DiffusionDB · Georgia Tech / Polo Club

    14M Stable Diffusion images with the exact prompts that produced them.

  • LLaVA-Instruct-150K · LLaVA Team (Liu et al.)

    158K GPT-4-generated visual instructions over COCO images. Research-only.