Modelplane Modelplane docs

All recipes

This document is for an unreleased version of Modelplane.

This document applies to the Modelplane main branch and not to the latest release v0.3.

Every validated recipe in one table: model, size, architecture, precision, and the verified hardware. Select a row for the full recipe.

ModelSizeArchPrecisionVerified onNotes
Qwen3-8B qwen8BDenseBF16EKS L4An 8.2B dense chat model on a single NVIDIA L4.
Qwen2.5-7B qwen7BDenseAWQ INT4Vultr A16A 7B dense chat model (AWQ INT4) on a single NVIDIA A16 on Vultr.
Qwen3-Coder-480B qwen480B A35BMoEBF16 / FP8EKS H200A 480B code MoE, multi-node BF16 over EFA or single-node FP8 on SGLang.
Qwen2.5-72B qwen72BDenseAWQ INT4AKSNebius A100H100A 72B dense chat model (AWQ INT4) on a single 80 GB GPU, on AKS and Nebius.
Kimi-K2 moonshotai1T A32BMoEINT4EKS H200A 1T MoE served prefill/decode disaggregated across two H200 nodes.
Llama-3.1-8B meta-llama8BDenseBF16EKSGKE L4An 8B dense chat model on a single NVIDIA L4.
GLM-4.5-Air zai-org106B A12BMoEGGUF IQ4_XSGKE A100A 106B MoE served from a GGUF checkpoint via llama.cpp on a single A100.
Nemotron-3.5-Lightning nvidia30B A3BMoENVFP4Nebius H100An open 30B MoE with 3B active parameters served NVFP4 on a single H100 on Nebius.
Laguna-S-2.1 poolside118B A8BMoEFP8Nebius H100A 118B code MoE served FP8 on a single 8x H100 node on Nebius.