Models
A catalogue of 89 open models you can run on the AETRON network: LLMs, image and video generation, speech and embeddings. Licence, memory requirements and parameters on every entry.
- Qwen2.5-0.5B-Instruct · Alibaba / Qwen
Tiny 0.5B instruction-tuned LLM — runs on edge devices and CPUs.
- Qwen2.5-0.5B-Instruct (bnb-4bit) · Alibaba / Qwen
4-bit (bnb NF4) build: Tiny 0.5B instruction-tuned LLM — runs on edge devices and CPUs.
- Qwen3.5-9B · Alibaba / Qwen
9B dense LLM from the Qwen3.5 generation (02.2026). Strong reasoning, coding and multilingual chat.
- Qwen3.5-9B (FP8-dynamic) · Alibaba / Qwen
Ready-made FP8 build (compressed-tensors): 9B dense LLM from the Qwen3.5 generation (02.2026). Strong reasoning, coding and multilingual chat.
- Qwen3.5-9B (w4a16) · Alibaba / Qwen
Ready-made INT4 W4A16 build (compressed-tensors): 9B dense LLM from the Qwen3.5 generation (02.2026). Strong reasoning, coding and multilingual chat.
- Qwen-Image · Alibaba Qwen
20B MMDiT text-to-image. Best-in-class text rendering, English and Chinese.
- Qwen3.6-27B · Alibaba / Qwen
Latest dense Qwen: natively multimodal 27B flagship.
- Qwen3.6-27B (NVFP4 (Blackwell)) · Alibaba / Qwen
Ready-made NVFP4 (Blackwell) build of Qwen3.6-27B. ModelOpt export - pending runtime support (transformers has no modelopt loader); catalog-only for now.
- Qwen3.6-27B (official FP8) · Alibaba / Qwen
Ready-made official FP8 build of Qwen3.6-27B.
- Qwen3.6-35B-A3B · Alibaba / Qwen
MoE 35B with ~3B active: near-27B quality, much faster.
- Qwen3.6-35B-A3B (NVFP4 (compressed-tensors)) · Alibaba / Qwen
Ready-made NVFP4 (compressed-tensors) build - runs on the fleet stack (compressed-tensors). Blackwell window.
- Qwen3.6-35B-A3B (official FP8) · Alibaba / Qwen
Ready-made official FP8 build of Qwen3.6-35B-A3B.
- Qwen3.5-4B · Alibaba / Qwen
Compact 4B for consumer GPUs and edge.
- Qwen3.5-4B (INT4 W4A16) · Alibaba / Qwen
Ready-made INT4 W4A16 build of Qwen3.5-4B.
- Qwen3.5-4B (FP8-dynamic) · Alibaba / Qwen
Ready-made FP8-dynamic build of Qwen3.5-4B.
- Qwen3.5-27B · Alibaba / Qwen
Previous-gen dense 27B, battle-tested.
- Qwen3.5-27B (official GPTQ-Int4) · Alibaba / Qwen
Ready-made official GPTQ-Int4 build of Qwen3.5-27B.
- Qwen3.5-9B (GGUF Q4_K_M) · Alibaba / Qwen (GGUF: unsloth)
llama.cpp GGUF build (Q4_K_M) of Qwen3.5-9B - one artifact runs on Metal, CUDA and CPU; the build the GGUF consensus engine was proven on.
- Qwen3.5-27B (GGUF Q4_K_M) · Alibaba / Qwen (GGUF: unsloth)
llama.cpp GGUF build (Q4_K_M) of Qwen3.5-27B - one artifact runs on Metal, CUDA and CPU.
- Qwen3.5-27B (GGUF Q8_0) · Alibaba / Qwen (GGUF: unsloth)
llama.cpp GGUF build (Q8_0, 8-bit) of Qwen3.5-27B - same model, higher-fidelity quant; a separate entry because a different artifact is a different model_hash.
- Qwen3.8-27B (GGUF UD-Q4_K_M) · Alibaba / Qwen (GGUF: unsloth)
llama.cpp GGUF build (Unsloth Dynamic Q4_K_M) of Qwen3.8-27B - one artifact runs on Metal, CUDA and CPU. 262144-token context. The upstream model is multimodal (image-text-to-text); this pin covers the TEXT weights only - the vision projector (mmproj) is a separate artifact and is not part of this manifest.
- Qwen3.5-27B (official FP8) · Alibaba / Qwen
Ready-made official FP8 build of Qwen3.5-27B.
- Qwen3.5-35B-A3B · Alibaba / Qwen
Qwen3.5 MoE 35B-A3B.
- Qwen3.5-35B-A3B (official GPTQ-Int4) · Alibaba / Qwen
Ready-made official GPTQ-Int4 build of Qwen3.5-35B-A3B.
- Qwen3.5-35B-A3B (official FP8) · Alibaba / Qwen
Ready-made official FP8 build of Qwen3.5-35B-A3B.
- Qwen3-Coder-30B-A3B · Alibaba / Qwen
Coding MoE with 256K context, ~3B active.
- Qwen3-Coder-30B-A3B (official FP8) · Alibaba / Qwen
Ready-made official FP8 build of Qwen3-Coder-30B-A3B.
- Gemma 4 E4B IT · Google
Gemma 4 efficient 4B instruction-tuned.
- Gemma 4 12B · Google
Mid-size Gemma 4.
- Gemma 4 31B · Google
Gemma 4 dense flagship.
- Gemma 4 31B (official QAT W4A16) · Google
Ready-made official QAT W4A16 build of Gemma 4 31B.
- Gemma 4 31B (FP8-block) · Google
Ready-made FP8-block build of Gemma 4 31B.
- Ministral 3 8B · Mistral AI
Vision-capable 8B from the edge-focused Ministral 3 line.
- Ministral 3 14B · Mistral AI
Larger Ministral 3 with vision.
- Ministral 3 14B (FP8-dynamic) · Mistral AI
Ready-made FP8-dynamic build of Ministral 3 14B.
- Phi-4 · Microsoft
Microsoft Phi-4 14.7B.
- Phi-4 (FP8-dynamic) · Microsoft
Ready-made FP8-dynamic build of Phi-4.
- Phi-4-mini · Microsoft
Phi-4-mini 3.8B for small GPUs and CPU.
- SmolLM3-3B · Hugging Face
Tiny 3B for edge and browser-adjacent uses.
- Nemotron-3 Nano 4B · NVIDIA
Hybrid Mamba-2 Nano 4B, long context.
- Nemotron-3 Nano 30B-A3B · NVIDIA
Nano MoE 30B-A3B, up to 1M context.
- Nemotron-3 Nano 30B-A3B (FP8) · NVIDIA
Ready-made FP8 build of Nemotron-3 Nano 30B-A3B.
- gpt-oss-20b · OpenAI
Native-MXFP4 21B MoE; o3-mini-class in 16 GB.
- GLM-4.7-Flash · Z.ai
Consumer GLM 31B.
- GLM-4.7-Flash (FP8-dynamic) · Z.ai
Ready-made FP8-dynamic build of GLM-4.7-Flash.
- DeepSeek-V4-Pro · DeepSeek
1.6T MoE frontier with built-in Think modes.
- DeepSeek-V4-Pro (NVFP4-FP8 (compressed-tensors)) · DeepSeek
Ready-made NVFP4-FP8 (compressed-tensors) build - runs on the fleet stack (compressed-tensors). Blackwell window.
- DeepSeek-V4-Flash · DeepSeek
Smaller, faster V4: 284B/13B active.
- DeepSeek-V4-Flash (NVFP4-FP8 (compressed-tensors)) · DeepSeek
Ready-made NVFP4-FP8 (compressed-tensors) build - runs on the fleet stack (compressed-tensors). Blackwell window.
- Kimi-K3 · Moonshot AI
2.8T MoE, first open 3T-class, MXFP4 QAT.
- Kimi-K3 (NVFP4) · Moonshot AI
Ready-made NVFP4 build of Kimi-K3.
- GLM-5.2 · Z.ai
753B MoE with sparse attention, 1M context.
- GLM-5.2 (official FP8) · Z.ai
Ready-made official FP8 build of GLM-5.2.
- Qwen3-Coder-Next · Alibaba / Qwen
SOTA open coder, 80B-A3B.
- Qwen3-Coder-Next (official FP8) · Alibaba / Qwen
Ready-made official FP8 build of Qwen3-Coder-Next.
- Llama 4 Scout · Meta
Llama 4 Scout 109B/17B active.
- Nemotron-3 Super 120B-A12B · NVIDIA
Super 120B-A12B: single 80 GB card at 4-bit.
- gpt-oss-120b · OpenAI
117B MoE in native MXFP4: one 80 GB GPU.
- Mistral Small 4 119B · Mistral AI
Unified instruct+reasoning 119B.
- Mistral Small 4 119B (official NVFP4) · Mistral AI
Ready-made official NVFP4 build of Mistral Small 4 119B. Mistral-format export (no config.json) - pending runtime support; catalog-only for now.
- FLUX.2 [klein] 4B · Black Forest Labs
Distilled consumer FLUX.2, 4B.
- FLUX.2 [klein] 9B · Black Forest Labs
Bigger klein: 9B.
- FLUX.2 [dev] · Black Forest Labs
32B DiT flagship.
- Qwen-Image 2.0 · Alibaba / Qwen
Qwen-Image 2.0: best-in-class text rendering.
- Z-Image Turbo · Tongyi-MAI
6B S3-DiT: best quality-per-GB, ~2s/image on 4090.
- Chroma1-HD · Lodestone Rock
Apache base for unrestricted finetuning.
- Wan 2.2 TI2V-5B · Alibaba / Wan-AI
5B text+image to video. Diffusers mirror: the upstream repo ships its VAE and T5 as pickle, which the loader rejects.
- Wan 2.2 T2V-A14B · Alibaba / Wan-AI
MoE video, 27B total / 14B active. Diffusers mirror; both experts stay resident, so VRAM is for the pair.
- HunyuanVideo 1.5 · Tencent
8.3B video model, 480p t2v. Diffusers mirror: the upstream repo needs Tencent's own hyvideo package and ships no text encoder.
- LTX-2.3 · Lightricks
Audio+video generation, speed champion. Manifest pins the distilled-1.1 checkpoint.
- Chatterbox · Resemble AI
Zero-shot voice cloning from ~10s. Manifest pins the multilingual v3 weights.
- Whisper large-v3-turbo · OpenAI
6x faster large-v3, 99 languages.
- Qwen3-VL-8B · Alibaba / Qwen
Dedicated VL line, 8B.
- InternVL3.5-8B · OpenGVLab
Strong on documents and PDFs. Ships auto_map (owner code), which the Level-1 weights-only runner refuses - catalog-only until a Level-2 path exists.
- SmolVLM2-2.2B · Hugging Face
Tiny VLM for edge.
- Chandra OCR 2 · Datalab
SOTA OCR on Qwen3.5, 40+ languages.
- olmOCR-2 7B · Allen AI
Fully open OCR for batch scale.
- Qwen3-Embedding-0.6B · Alibaba / Qwen
Text embedding leader line, CPU-capable size.
- Qwen3-Embedding-8B · Alibaba / Qwen
Top MMTEB text embeddings.
- EmbeddingGemma-300m · Google
On-device embeddings, 100+ languages.
- MiniMax H3 (Hailuo 3) · MiniMax
33B joint video+audio DiT (Hailuo 3). Root modular-diffusers layout, T2VA stack; needs a modular runtime the runner does not have yet.
- MiniMax H3 (GGUF Q4_K_M, sd.cpp) · MiniMax (GGUF: leejet / Comfy-Org)
Video+stereo audio generation on the sd.cpp engine: T2VA, start/end frame (I2VA/FL2VA). Q4 DiT + Qwen3-VL-32B TE; runs on one 24 GB GPU (needs ~48 GB host RAM).
- MiniMax H3 Ref2VA (GGUF Q4_K_M, sd.cpp) · MiniMax (GGUF: leejet / Comfy-Org)
MiniMax H3 reference-conditioned video (Ref2VA): up to 9 image, 3 video, 3 audio refs, cited as <Picture N>/<Video N>/<Audio N>.
- MiniMax Music 3 (GGUF Q8_0 DiT) · MiniMax (GGUF DiT: Abiray)
Two-stage text-to-music (8B LLM + flow-matching DiT): songs up to 5 min, 32 kHz stereo. Official components + Q8_0 GGUF DiT; awaits a diffusers release.
- MiniMax M2.7 · MiniMax
MiniMax's frontier MoE LLM (M2.7 line).
- MiniMax M2.7 (NVFP4) · MiniMax
Ready-made NVFP4 (Blackwell) build of MiniMax M2.7. ModelOpt export - pending runtime support (transformers has no modelopt loader); catalog-only for now.
- Kimi K2.7 Code · Moonshot AI
Moonshot's agentic-coding MoE (~300B, K2.7 line).
- Kimi K2.7 Code (NVFP4) · Moonshot AI
Ready-made NVFP4 (Blackwell) build of Kimi K2.7 Code. ModelOpt export - pending runtime support (transformers has no modelopt loader); catalog-only for now.
- ACE-Step v1 3.5B · ACE Studio / StepFun
Text-to-music diffusion model, 3.5B. Generates full songs with vocals from a prompt; 19 languages.