Models

A catalogue of 89 open models you can run on the AETRON network: LLMs, image and video generation, speech and embeddings. Licence, memory requirements and parameters on every entry.

  • Qwen2.5-0.5B-Instruct · Alibaba / Qwen

    Tiny 0.5B instruction-tuned LLM — runs on edge devices and CPUs.

  • Qwen2.5-0.5B-Instruct (bnb-4bit) · Alibaba / Qwen

    4-bit (bnb NF4) build: Tiny 0.5B instruction-tuned LLM — runs on edge devices and CPUs.

  • Qwen3.5-9B · Alibaba / Qwen

    9B dense LLM from the Qwen3.5 generation (02.2026). Strong reasoning, coding and multilingual chat.

  • Qwen3.5-9B (FP8-dynamic) · Alibaba / Qwen

    Ready-made FP8 build (compressed-tensors): 9B dense LLM from the Qwen3.5 generation (02.2026). Strong reasoning, coding and multilingual chat.

  • Qwen3.5-9B (w4a16) · Alibaba / Qwen

    Ready-made INT4 W4A16 build (compressed-tensors): 9B dense LLM from the Qwen3.5 generation (02.2026). Strong reasoning, coding and multilingual chat.

  • Qwen-Image · Alibaba Qwen

    20B MMDiT text-to-image. Best-in-class text rendering, English and Chinese.

  • Qwen3.6-27B · Alibaba / Qwen

    Latest dense Qwen: natively multimodal 27B flagship.

  • Qwen3.6-27B (NVFP4 (Blackwell)) · Alibaba / Qwen

    Ready-made NVFP4 (Blackwell) build of Qwen3.6-27B. ModelOpt export - pending runtime support (transformers has no modelopt loader); catalog-only for now.

  • Qwen3.6-27B (official FP8) · Alibaba / Qwen

    Ready-made official FP8 build of Qwen3.6-27B.

  • Qwen3.6-35B-A3B · Alibaba / Qwen

    MoE 35B with ~3B active: near-27B quality, much faster.

  • Qwen3.6-35B-A3B (NVFP4 (compressed-tensors)) · Alibaba / Qwen

    Ready-made NVFP4 (compressed-tensors) build - runs on the fleet stack (compressed-tensors). Blackwell window.

  • Qwen3.6-35B-A3B (official FP8) · Alibaba / Qwen

    Ready-made official FP8 build of Qwen3.6-35B-A3B.

  • Qwen3.5-4B · Alibaba / Qwen

    Compact 4B for consumer GPUs and edge.

  • Qwen3.5-4B (INT4 W4A16) · Alibaba / Qwen

    Ready-made INT4 W4A16 build of Qwen3.5-4B.

  • Qwen3.5-4B (FP8-dynamic) · Alibaba / Qwen

    Ready-made FP8-dynamic build of Qwen3.5-4B.

  • Qwen3.5-27B · Alibaba / Qwen

    Previous-gen dense 27B, battle-tested.

  • Qwen3.5-27B (official GPTQ-Int4) · Alibaba / Qwen

    Ready-made official GPTQ-Int4 build of Qwen3.5-27B.

  • Qwen3.5-9B (GGUF Q4_K_M) · Alibaba / Qwen (GGUF: unsloth)

    llama.cpp GGUF build (Q4_K_M) of Qwen3.5-9B - one artifact runs on Metal, CUDA and CPU; the build the GGUF consensus engine was proven on.

  • Qwen3.5-27B (GGUF Q4_K_M) · Alibaba / Qwen (GGUF: unsloth)

    llama.cpp GGUF build (Q4_K_M) of Qwen3.5-27B - one artifact runs on Metal, CUDA and CPU.

  • Qwen3.5-27B (GGUF Q8_0) · Alibaba / Qwen (GGUF: unsloth)

    llama.cpp GGUF build (Q8_0, 8-bit) of Qwen3.5-27B - same model, higher-fidelity quant; a separate entry because a different artifact is a different model_hash.

  • Qwen3.8-27B (GGUF UD-Q4_K_M) · Alibaba / Qwen (GGUF: unsloth)

    llama.cpp GGUF build (Unsloth Dynamic Q4_K_M) of Qwen3.8-27B - one artifact runs on Metal, CUDA and CPU. 262144-token context. The upstream model is multimodal (image-text-to-text); this pin covers the TEXT weights only - the vision projector (mmproj) is a separate artifact and is not part of this manifest.

  • Qwen3.5-27B (official FP8) · Alibaba / Qwen

    Ready-made official FP8 build of Qwen3.5-27B.

  • Qwen3.5-35B-A3B · Alibaba / Qwen

    Qwen3.5 MoE 35B-A3B.

  • Qwen3.5-35B-A3B (official GPTQ-Int4) · Alibaba / Qwen

    Ready-made official GPTQ-Int4 build of Qwen3.5-35B-A3B.

  • Qwen3.5-35B-A3B (official FP8) · Alibaba / Qwen

    Ready-made official FP8 build of Qwen3.5-35B-A3B.

  • Qwen3-Coder-30B-A3B · Alibaba / Qwen

    Coding MoE with 256K context, ~3B active.

  • Qwen3-Coder-30B-A3B (official FP8) · Alibaba / Qwen

    Ready-made official FP8 build of Qwen3-Coder-30B-A3B.

  • Gemma 4 E4B IT · Google

    Gemma 4 efficient 4B instruction-tuned.

  • Gemma 4 12B · Google

    Mid-size Gemma 4.

  • Gemma 4 31B · Google

    Gemma 4 dense flagship.

  • Gemma 4 31B (official QAT W4A16) · Google

    Ready-made official QAT W4A16 build of Gemma 4 31B.

  • Gemma 4 31B (FP8-block) · Google

    Ready-made FP8-block build of Gemma 4 31B.

  • Ministral 3 8B · Mistral AI

    Vision-capable 8B from the edge-focused Ministral 3 line.

  • Ministral 3 14B · Mistral AI

    Larger Ministral 3 with vision.

  • Ministral 3 14B (FP8-dynamic) · Mistral AI

    Ready-made FP8-dynamic build of Ministral 3 14B.

  • Phi-4 · Microsoft

    Microsoft Phi-4 14.7B.

  • Phi-4 (FP8-dynamic) · Microsoft

    Ready-made FP8-dynamic build of Phi-4.

  • Phi-4-mini · Microsoft

    Phi-4-mini 3.8B for small GPUs and CPU.

  • SmolLM3-3B · Hugging Face

    Tiny 3B for edge and browser-adjacent uses.

  • Nemotron-3 Nano 4B · NVIDIA

    Hybrid Mamba-2 Nano 4B, long context.

  • Nemotron-3 Nano 30B-A3B · NVIDIA

    Nano MoE 30B-A3B, up to 1M context.

  • Nemotron-3 Nano 30B-A3B (FP8) · NVIDIA

    Ready-made FP8 build of Nemotron-3 Nano 30B-A3B.

  • gpt-oss-20b · OpenAI

    Native-MXFP4 21B MoE; o3-mini-class in 16 GB.

  • GLM-4.7-Flash · Z.ai

    Consumer GLM 31B.

  • GLM-4.7-Flash (FP8-dynamic) · Z.ai

    Ready-made FP8-dynamic build of GLM-4.7-Flash.

  • DeepSeek-V4-Pro · DeepSeek

    1.6T MoE frontier with built-in Think modes.

  • DeepSeek-V4-Pro (NVFP4-FP8 (compressed-tensors)) · DeepSeek

    Ready-made NVFP4-FP8 (compressed-tensors) build - runs on the fleet stack (compressed-tensors). Blackwell window.

  • DeepSeek-V4-Flash · DeepSeek

    Smaller, faster V4: 284B/13B active.

  • DeepSeek-V4-Flash (NVFP4-FP8 (compressed-tensors)) · DeepSeek

    Ready-made NVFP4-FP8 (compressed-tensors) build - runs on the fleet stack (compressed-tensors). Blackwell window.

  • Kimi-K3 · Moonshot AI

    2.8T MoE, first open 3T-class, MXFP4 QAT.

  • Kimi-K3 (NVFP4) · Moonshot AI

    Ready-made NVFP4 build of Kimi-K3.

  • GLM-5.2 · Z.ai

    753B MoE with sparse attention, 1M context.

  • GLM-5.2 (official FP8) · Z.ai

    Ready-made official FP8 build of GLM-5.2.

  • Qwen3-Coder-Next · Alibaba / Qwen

    SOTA open coder, 80B-A3B.

  • Qwen3-Coder-Next (official FP8) · Alibaba / Qwen

    Ready-made official FP8 build of Qwen3-Coder-Next.

  • Llama 4 Scout · Meta

    Llama 4 Scout 109B/17B active.

  • Nemotron-3 Super 120B-A12B · NVIDIA

    Super 120B-A12B: single 80 GB card at 4-bit.

  • gpt-oss-120b · OpenAI

    117B MoE in native MXFP4: one 80 GB GPU.

  • Mistral Small 4 119B · Mistral AI

    Unified instruct+reasoning 119B.

  • Mistral Small 4 119B (official NVFP4) · Mistral AI

    Ready-made official NVFP4 build of Mistral Small 4 119B. Mistral-format export (no config.json) - pending runtime support; catalog-only for now.

  • FLUX.2 [klein] 4B · Black Forest Labs

    Distilled consumer FLUX.2, 4B.

  • FLUX.2 [klein] 9B · Black Forest Labs

    Bigger klein: 9B.

  • FLUX.2 [dev] · Black Forest Labs

    32B DiT flagship.

  • Qwen-Image 2.0 · Alibaba / Qwen

    Qwen-Image 2.0: best-in-class text rendering.

  • Z-Image Turbo · Tongyi-MAI

    6B S3-DiT: best quality-per-GB, ~2s/image on 4090.

  • Chroma1-HD · Lodestone Rock

    Apache base for unrestricted finetuning.

  • Wan 2.2 TI2V-5B · Alibaba / Wan-AI

    5B text+image to video. Diffusers mirror: the upstream repo ships its VAE and T5 as pickle, which the loader rejects.

  • Wan 2.2 T2V-A14B · Alibaba / Wan-AI

    MoE video, 27B total / 14B active. Diffusers mirror; both experts stay resident, so VRAM is for the pair.

  • HunyuanVideo 1.5 · Tencent

    8.3B video model, 480p t2v. Diffusers mirror: the upstream repo needs Tencent's own hyvideo package and ships no text encoder.

  • LTX-2.3 · Lightricks

    Audio+video generation, speed champion. Manifest pins the distilled-1.1 checkpoint.

  • Chatterbox · Resemble AI

    Zero-shot voice cloning from ~10s. Manifest pins the multilingual v3 weights.

  • Whisper large-v3-turbo · OpenAI

    6x faster large-v3, 99 languages.

  • Qwen3-VL-8B · Alibaba / Qwen

    Dedicated VL line, 8B.

  • InternVL3.5-8B · OpenGVLab

    Strong on documents and PDFs. Ships auto_map (owner code), which the Level-1 weights-only runner refuses - catalog-only until a Level-2 path exists.

  • SmolVLM2-2.2B · Hugging Face

    Tiny VLM for edge.

  • Chandra OCR 2 · Datalab

    SOTA OCR on Qwen3.5, 40+ languages.

  • olmOCR-2 7B · Allen AI

    Fully open OCR for batch scale.

  • Qwen3-Embedding-0.6B · Alibaba / Qwen

    Text embedding leader line, CPU-capable size.

  • Qwen3-Embedding-8B · Alibaba / Qwen

    Top MMTEB text embeddings.

  • EmbeddingGemma-300m · Google

    On-device embeddings, 100+ languages.

  • MiniMax H3 (Hailuo 3) · MiniMax

    33B joint video+audio DiT (Hailuo 3). Root modular-diffusers layout, T2VA stack; needs a modular runtime the runner does not have yet.

  • MiniMax H3 (GGUF Q4_K_M, sd.cpp) · MiniMax (GGUF: leejet / Comfy-Org)

    Video+stereo audio generation on the sd.cpp engine: T2VA, start/end frame (I2VA/FL2VA). Q4 DiT + Qwen3-VL-32B TE; runs on one 24 GB GPU (needs ~48 GB host RAM).

  • MiniMax H3 Ref2VA (GGUF Q4_K_M, sd.cpp) · MiniMax (GGUF: leejet / Comfy-Org)

    MiniMax H3 reference-conditioned video (Ref2VA): up to 9 image, 3 video, 3 audio refs, cited as <Picture N>/<Video N>/<Audio N>.

  • MiniMax Music 3 (GGUF Q8_0 DiT) · MiniMax (GGUF DiT: Abiray)

    Two-stage text-to-music (8B LLM + flow-matching DiT): songs up to 5 min, 32 kHz stereo. Official components + Q8_0 GGUF DiT; awaits a diffusers release.

  • MiniMax M2.7 · MiniMax

    MiniMax's frontier MoE LLM (M2.7 line).

  • MiniMax M2.7 (NVFP4) · MiniMax

    Ready-made NVFP4 (Blackwell) build of MiniMax M2.7. ModelOpt export - pending runtime support (transformers has no modelopt loader); catalog-only for now.

  • Kimi K2.7 Code · Moonshot AI

    Moonshot's agentic-coding MoE (~300B, K2.7 line).

  • Kimi K2.7 Code (NVFP4) · Moonshot AI

    Ready-made NVFP4 (Blackwell) build of Kimi K2.7 Code. ModelOpt export - pending runtime support (transformers has no modelopt loader); catalog-only for now.

  • ACE-Step v1 3.5B · ACE Studio / StepFun

    Text-to-music diffusion model, 3.5B. Generates full songs with vocals from a prompt; 19 languages.