$ ls ~/recipes/

Recipes

How I serve models on DGX Spark. Every config here is real and in use: the lmswitch recipes I switch between day to day, and the reproducible recipe behind every run on the DGX Spark bench.

$ ls ~/ai-models/*.yaml

lmswitch configs from jvr0x/ai-models (llama.cpp, vLLM, dual-Spark vLLM and ExLlamaV3 jobs). Each links to its YAML on GitHub.

DeepSeek-V4-Flash 0731 Abliterated 1M [DUAL 2x SPARK]

Keys abliterated DeepSeek-V4-Flash-0731 - TP=2 across spark + gigabyte over CX7 (1M ctx).

deepseek-v4-flash-0731-ablit-dual.yaml · vllm-dual · drowzeys/keys-DeepSeekV4-Flash-GA-0731-Dspark-Abliterated-32-32

DeepSeek-V4-Flash 0731 1M [DUAL 2x SPARK]

DeepSeek V4 Flash 0731 - TP=2 across spark + gigabyte over CX7 (1M ctx).

deepseek-v4-flash-0731-dual.yaml · vllm-dual · deepseek-ai/DeepSeek-V4-Flash-0731

DeepSeek-V4-Flash DSpark 1M [DUAL 2x SPARK]

DeepSeek V4 Flash DSpark - TP=2 across spark + gigabyte over CX7 (1M ctx).

deepseek-v4-flash-dspark.yaml · vllm-dual · deepseek-ai/DeepSeek-V4-Flash-DSpark

DiffusionGemma-26B-IT

diffusiongemma-26b.yaml · vllm · nvidia/diffusiongemma-26B-A4B-it-NVFP4

Gemma-4-12B-IT NVFP4 (Unsloth)

Gemma-4-12B-IT NVFP4 (Unsloth) - small-footprint vLLM profile.

gemma-4-12b-it-nvfp4.yaml · vllm · unsloth/gemma-4-12b-it-NVFP4

Gemma-4-12B-IT QAT Q4 MTP

Gemma-4-12B-IT QAT GGUF (Unsloth) - vision + MTP speculative decoding.

gemma-4-12b-it-qat-mtp.yaml · llama · unsloth/gemma-4-12B-it-qat-GGUF/gemma-4-12B-it-qat-UD-Q4_K_XL.gguf

Gemma-4-12B Instruct

gemma-4-12b-it.yaml · llama · unsloth/gemma-4-12b-it-GGUF/gemma-4-12b-it-Q4_K_M.gguf

Gemma-4-26B-A4B Instruct

gemma-4-26b-a4b-it.yaml · vllm · google/gemma-4-26B-A4B-it

Gemma-4-31B-IT

gemma4-31b.yaml · vllm · nvidia/Gemma-4-31B-IT-NVFP4

Gemma-4-12B Coder

gemma4-coder.yaml · llama · yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF/gemma4-coding-Q8_0.gguf

GPT-OSS-20B

gpt-oss-20b.yaml · llama · openai/gpt-oss-20b-GGUF/gpt-oss-20b-UD-Q4_K_XL.gguf

Kimi-Linear-48B Q8

kimi-linear.yaml · llama · bartowski/moonshotai_Kimi-Linear-48B-A3B-Instruct-GGUF/moonshotai_Kimi-Linear-48B-A3B-Instruct-Q8_0/moonshotai_Kimi-Linear-48B-A3B-Instruct-Q8_0-00001-of-00002.gguf

Mistral-Medium-3.5-128B NVFP4 [BENCH]

lmswitch BENCH PROFILE of Mistral-Medium-3.5-128B NVFP4 - dense 128B via vLLM.

mistral-medium-3.5-128b-nvfp4-bench.yaml · vllm · nvidia/Mistral-Medium-3.5-128B-NVFP4

Nemotron-3-Super-120B NVFP4 [BENCH]

lmswitch BENCH PROFILE of Nemotron-3-Super-120B-A12B NVFP4 via vLLM (stock image).

nemotron-3-super-nvfp4-vllm.yaml · vllm · nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4

Nemotron-3-Super-120B NVFP4

lmswitch recipe for NVIDIA Nemotron-3-Super-120B-A12B-NVFP4 on 1× DGX Spark.

nemotron-3-super-nvfp4.yaml · vllm · nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4

Nemotron-Nano-30B

nemotron.yaml · llama · nemotron3-gguf/Nemotron-3-Nano-30B-A3B-UD-Q8_K_XL.gguf

Nemotron-3-Nano-30B FP8

nemotron3-nano-fp8.yaml · vllm · nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-FP8

Nemotron-3-Nano Omni NVFP4

nemotron3-nano-omni-nvfp4.yaml · vllm · nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-NVFP4

Nex-N2-Pro 397B-A17B (IQ1_M)

nex-n2-pro.yaml · llama · Nex-N2-Pro-IQ1_M/IQ1_M/Nex-397B-A17B-IQ1_M-00001-of-00005.gguf

Nemotron-Labs-3-Puzzle-75B-A9B NVFP4

lmswitch recipe for NVIDIA Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4 on 1× DGX Spark.

nvidia-nemotron-labs-3-puzzle-75b.yaml · vllm · nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4

Ornith-1.0-35B NVFP4 (AEON) [BENCH]

lmswitch BENCH PROFILE of Ornith-1.0-35B-AEON NVFP4 - hybrid Mamba/Attention.

ornith-35b-nvfp4-aeon7.yaml · vllm · AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4

Ornith-1.0-35B Q8

ornith-35b-q8.yaml · llama · ornith/Ornith-1.0-35B-Q8_0/ornith-1.0-35b-Q8_0.gguf

Qwen-AgentWorld-35B-A3B

qwen-agentworld-stock.yaml · vllm · Qwen-AgentWorld-35B-A3B

Qwen3-Coder-Next Q4

qwen3-coder-next.yaml · llama · unsloth/Qwen3-Coder-Next-GGUF/Qwen3-Coder-Next-UD-Q4_K_XL.gguf

Qwen3-VL-30B-A3B Instruct

qwen3-vl-30b.yaml · llama · unsloth/Qwen3-VL-30B-A3B-Instruct-GGUF/Qwen3-VL-30B-A3B-Instruct-UD-Q4_K_XL.gguf

Qwen3-VL-8B Instruct Q4

qwen3-vl-8b.yaml · llama · unsloth/Qwen3-VL-8B-Instruct-GGUF/Qwen3-VL-8B-Instruct-UD-Q4_K_XL.gguf

Qwen3.6-27B AEON Uncensored NVFP4 +DFlash

qwen3.6-27b-nvfp4-aeon7.yaml · vllm · AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-NVFP4

Qwen3.6-27B NVFP4 (MiaAI-Lab repo config)

Faithful port of MiaAI-Lab/Qwen3.6-27B-NVFP4-vLLM start.sh, as a lmswitch recipe.

qwen3.6-27b-nvfp4-miaai.yaml · vllm · nvidia/Qwen3.6-27B-NVFP4

Qwen3.6-27B NVFP4 (Unsloth) [BENCH]

lmswitch BENCH PROFILE of qwen3.6-27b-nvfp4-unsloth - local NVFP4 dense hybrid VLM.

qwen3.6-27b-nvfp4-unsloth-bench.yaml · vllm · unsloth/Qwen3.6-27B-NVFP4

Qwen3.6-27B NVFP4 [UNSLOTH]

lmswitch recipe for Qwen3.6-27B-NVFP4 (unsloth) on 1× DGX Spark.

qwen3.6-27b-nvfp4-unsloth.yaml · vllm · unsloth/Qwen3.6-27B-NVFP4

Qwen3.6-35B-A3B NVFP4-Fast [UNSLOTH]

lmswitch recipe for Qwen3.6-35B-A3B-NVFP4-Fast (unsloth) on 1× DGX Spark.

qwen3.6-35b-a3b-nvfp4-fast.yaml · vllm · unsloth/Qwen3.6-35B-A3B-NVFP4-Fast

Qwen3.6-35B heretic NVFP4 +DFlash

qwen3.6-35b-heretic-dflash-aeon7.yaml · vllm · AEON-7/Qwen3.6-35B-A3B-heretic-NVFP4

Qwen3.6-35B-A3B NVFP4 Fast (Unsloth) [BENCH]

lmswitch BENCH PROFILE of qwen3.6-35b-nvfp4-unsloth-fast - local NVFP4 MoE hybrid.

qwen3.6-35b-nvfp4-fast.yaml · vllm · unsloth/Qwen3.6-35B-A3B-NVFP4-Fast

Qwen3.6-35B-A3B NVFP4 (NVIDIA)

qwen3.6-35b-nvfp4-nvidia-aeon7.yaml · vllm · nvidia/qwen3.6-35b-a3b-nvfp4

Qwen3.6-35B-A3B Q8 MTP

qwen3.6-35b-q8-mtp.yaml · llama · unsloth/Qwen3.6-35B-A3B-MTP-GGUF/Qwen3.6-35B-A3B-UD-Q8_K_XL.gguf

Qwen3.6-35B Distilled Q8

qwen3.6-distill-q8.yaml · llama · lordx64/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-Q8_0-GGUF/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled.Q8_0.gguf

Qwen3.6-35B Distilled

qwen3.6-distill.yaml · llama · lordx64/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-IQ4_XS-GGUF/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled.IQ4_XS.gguf

Qwen3.6-35B Kimi-K2.6 Distilled Q8

qwen3.6-kimi-distill.yaml · llama · lordx64/Qwen3.6-35B-A3B-Kimi-K2.6-Reasoning-Distilled-GGUF/Qwen3.6-35B-A3B-Kimi-K2.6-Reasoning-Distilled.Q8_0.gguf

Qwopus3.6-27B Coder MTP (Q8_0)

qwopus3.6-coder.yaml · llama · Jackrong/Qwopus3.6-27B-Coder-MTP-GGUF/Qwopus3.6-27B-Coder-MTP-Q8_0.gguf

RED-SNOW-5.3-FLASH EXL3-SAGE 4.91bpw [DUAL 2x SPARK job]

lmswitch job for RED-SNOW-5.3-FLASH EXL3 SAGE 4.91bpw across spark + gigabyte.

red-snow-exl3-sage-dual.yaml · exllamav3-dual · vcruz305/RED-SNOW-5.3-FLASH-EXL3-SAGE-4.91bpw

Step-3.7-Flash (IQ4_XS)

step-3.7-flash.yaml · llama · unsloth/Step-3.7-Flash-GGUF/UD-IQ4_XS/Step-3.7-Flash-UD-IQ4_XS-00001-of-00003.gguf

$ ls ~/forks/

DeepSeek-V4.1-Flash EXL3 on one DGX Spark

Serving DeepSeek-V4.1-Flash as EXL3 on a single DGX Spark, about 17.5 tok/s with the DSpark drafter.

fork · ExLlamaV3 · 1× DGX Spark

DeepSeekEXL3Speculative decoding

$ ls ~/dgx-spark-bench/recipes/

One entry per measured config on the bench dashboard, linked to the recipe that produced it. Where a bench config is the same file as an lmswitch recipe, it links to that too.

DeepSeek-V4-Flash 0731 1M [DUAL 2xSpark]

NVFP4-DS-MLA KV + b12x MoE + DSpark spec-decode (K=5) · vLLM 0.25.2.dev0+g752a3a504 (Anemll dspark-vllm-gx10 0.1.1, TP=2)

DeepSeek-V4-Flash DSpark 1M [DUAL 2xSpark]

NVFP4-DS-MLA KV + b12x MoE + DSpark spec-decode · vLLM 0.25 (Anemll dspark-vllm-gx10, TP=2)

DeepSeek-V4-Flash (0731)

Upstream recipe by MiaAI-Lab, measured on the bench.

EXL3 3.0 bpw (0xSero spark build) · sparkinfer (vLLM 26.02 base)

DeepSeek-V4-Flash-Spark

GGUF Q3 REAP · llama.cpp

Gemma-4-12B-IT

GGUF Q4_K_M · llama.cpp

Gemma-4-26B-A4B-IT

BF16 · vLLM 0.23

Gemma-4-31B-IT

NVFP4 · vLLM 0.23

GLM-5.3-Flash

Upstream recipe by vcruz305, measured on the bench.

EXL3 K2 (2-bit routed experts, ~91 GiB pack) · vLLM + vllm-exl3 0.3.1 (Cruz prebuilt venv)

GPT-OSS-20B

GGUF Q4_K_XL · llama.cpp

Kimi-Linear-48B-A3B

GGUF Q8_0 · llama.cpp

Kolibri-1

Upstream recipe by Aleph-Alpha, measured on the bench.

FP8 · vLLM

Laguna-S-2.1 [DUAL 2x SPARK]

NVFP4 + DFlash · vLLM (TP=2, native --nnodes 2)

Laguna-S-2.1 [nspec=15]

NVFP4 + DFlash (15 tok, untuned) · vLLM

Laguna-S-2.1

NVFP4 + DFlash · vLLM

MiMo-V2.5 Omni [DUAL 2xSpark]

NVFP4 + NVFP4-KV + MTP1 · vLLM (vllm-dual-ray, Ray TP=2)

MiniMax-M3 428B [DUAL 2xSpark]

W4A16 GPTQ + NVFP4-KV (b12x) + EAGLE3 · vLLM (vllm-dual, TP=2)

Nemotron-Labs-3-Puzzle-75B-A9B

NVFP4 + MTP · vLLM 0.23 (AEON)

Nemotron-3-Nano-30B-A3B

FP8 · vLLM 0.23

Nemotron-3-Nano-30B-A3B

GGUF Q8_K_XL · llama.cpp

Nemotron-3-Nano-Omni-30B-A3B

NVFP4 · vLLM 0.23

Nex-N2-Pro-397B-A17B

GGUF IQ1_M · llama.cpp

Ornith-1.0-35B-AEON

NVFP4-mixed (compressed-tensors) · vLLM 0.23 (AEON)

Qwen-AgentWorld-35B-A3B

BF16 · vLLM 0.23

Qwen3-Coder-Next-80B-A3B

GGUF Q4_K_XL · llama.cpp

Qwen3.5-397B-A17B [DUAL 2xSpark]

GGUF UD-IQ4_NL · llama.cpp (RPC, 2 nodes)

Qwen3.6-27B-AEON

NVFP4-mixed (compressed-tensors) · vLLM 0.23

Qwen3.6-27B

NVFP4 · vLLM 0.23

Qwen3.6-27B-NVFP4

NVFP4-mixed (dense hybrid) · vLLM 0.23 (AEON)

Qwen3.6-27B-NVFP4 (Unsloth)

NVFP4-mixed (dense hybrid) · vLLM 0.23 (AEON)

Qwen3.6-35B-heretic

NVFP4-mixed (compressed-tensors) · vLLM 0.23

Qwen3.6-35B-A3B

NVFP4-mixed (spec OFF) · vLLM 0.23 (AEON)

Qwen3.6-35B-A3B

NVFP4-mixed +MTP · vLLM 0.23 (AEON)

Qwen3.6-35B-A3B

GGUF Q4_K_XL · llama.cpp

Qwen3.6-35B-A3B

GGUF Q8_K_XL · llama.cpp

Qwen3.8-27B

NVFP4-mixed (ModelOpt) · SGLang nightly-dev-20260814-c4271c3f

Qwen3.8-Flash-Next

Upstream recipe by MiaAI-Lab, measured on the bench.

NVFP4 (Mia-AiLab) · vLLM qwen38-flash-next image

Qwopus3.6-27B-Coder

GGUF Q8_0 · llama.cpp

Step-3.7-Flash

GGUF IQ4_XS · llama.cpp

SuperQwen3.8-27B-abliterated

NVFP4 W4A4 group-16 (ModelOpt) · vLLM (anemll dspark-vllm-gx10 0.1.1)