# K2 Horizon MoVA 36B A4B

URL: https://interfaze.ai/models/ifmk2-horizon-mova-36b-a4b

[All models](https://interfaze.ai/models)

K2 Horizon MoVA 36B A4B by IFM, a text-generation model. Understand and compare features, benchmarks, and capabilities.

## Comparison

| Feature | K2 Horizon MoVA 36B A4B | Interfaze |
| --- | --- | --- |
| Input Modalities | text | image, text, audio, video, document |
| Native OCR | No | Yes |
| Long Document Processing | No | Yes |
| Language Support | unknown | 162+ |
| Native Speech-to-Text | No | Yes |
| Native Object Detection | No | Yes |
| Guardrail Controls | No | Yes |
| Context Input Size | 524.3K | 1M |
| Tool Calling | Yes | Tool calling supported + built in browser, code execution and web search |

### Scaling

| Feature | K2 Horizon MoVA 36B A4B | Interfaze |
| --- | --- | --- |
| Scaling | Self-hosted/Provider-hosted with quantization | Unlimited |

[Try Interfaze](https://interfaze.ai/dashboard)[Read the Docs](https://interfaze.ai/docs)

View model card on [Hugging Face](https://huggingface.co/IFM/K2-Horizon-MoVA-36B-A4B)

K2-Horizon-MoVA-36B-A4B is the sparse member of the K2-Horizon family: a Mixture-of-Experts model with Mixture-of-Values attention (MoVA) that stores 36B parameters and runs 4B per token. We have released the final checkpoint; intermediate checkpoints, along with the data and the training code, will be released.

## K2-Horizon-MoVA-36B-A4B Highlights

-   **Frontier-class results at 4B active parameters.** On agentic and reasoning benchmarks it outscores open weight dense (approximately 30B model size) and MoE models up to 15× its size; and also performs competitively against closed frontier models (see [Benchmark Results](https://interfaze.ai/models/ifmk2-horizon-mova-36b-a4b#benchmark-results)).
-   **512K context.** Native 524,288-token context from the midtraining stages onward.
-   **Intermediate checkpoints.** Intermediate checkpoints will be released so capability changes can be studied across training rather than at a single checkpoint.
-   **Fully open.** Training data/recipe and the training code will be made public.

## Benchmark Results

## Quickstart

### Serving

vLLM, recipe at [recipes.vllm.ai/IFM](https://recipes.vllm.ai/IFM):

```
vllm serve IFM/K2-Horizon-MoVA-36B-A4B \
  --revision main \
  --tensor-parallel-size 2 \
  --enable-expert-parallel \
  --trust-remote-code \
  --dtype bfloat16 \
  --max-model-len 131072 \
  --reasoning-parser k2_horizon \
  --tool-call-parser k2_horizon \
  --enable-auto-tool-choice
```

SGLang recipe validated on 2× H200 in the [SGLang K2 Horizon cookbook](https://docs.sglang.io/cookbook/autoregressive/IFM/K2-Horizon):

```
python3 -m sglang.launch_server \
  --model-path IFM/K2-Horizon-MoVA-36B-A4B \
  --revision main \
  --tp 2 \
  --ep 2 \
  --dtype bfloat16 \
  --attention-backend fa3 \
  --json-model-override-args '{"xllm_source_router_gemm_partitions":2}' \
  --reasoning-parser k2_horizon \
  --tool-call-parser k2_horizon \
  --host 0.0.0.0 --port 30000
```

### API Usage

> \[!Tip\] Recommended settings: `reasoning_effort="high"`, `temperature=1.0`, `top_p=0.95`. Reasoning depth is selected per request through `chat_template_kwargs`. Thinking is returned in `reasoning_content` and the answer in `content`.

```
from openai import OpenAI

client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
response = client.chat.completions.create(
    model="IFM/K2-Horizon-MoVA-36B-A4B",
    messages=[{"role": "user", "content": "Explain the result step by step."}],
    temperature=1.0,
    top_p=0.95,
    max_tokens=32768,
    extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}},
)
message = response.choices[0].message
print("Reasoning:", getattr(message, "reasoning_content", None))
print("Answer:", message.content)
```

### Transformers

Validated with Transformers 5.15.0, PyTorch 2.13.0, Safetensors 0.8.0.

```
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "IFM/K2-Horizon-MoVA-36B-A4B"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id, device_map="auto", dtype="bfloat16", low_cpu_mem_usage=True, trust_remote_code=True
)

inputs = tokenizer("Explain why long-context evaluation is difficult.", return_tensors="pt").to(model.device)
inputs.pop("token_type_ids", None)
outputs = model.generate(**inputs, max_new_tokens=32768, temperature=1.0, top_p=0.95, do_sample=True)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```

## Best Practices

1.  **Reasoning effort: always `high`.** All reported results use high reasoning effort. Pass `{"chat_template_kwargs": {"reasoning_effort": "high"}}` on every request.
2.  **Sampling parameters.** `temperature=1.0`, `top_p=0.95`.
3.  **Serving.** Use the validated SGLang recipe above: BF16, TP=2, FlashAttention-3, and the `xllm_source_router_gemm_partitions` override, which preserves the checkpoint's router numerics. Full recipes for every K2-Horizon size, with measured H200 latency and throughput, are in the [SGLang cookbook](https://docs.sglang.io/cookbook/autoregressive/IFM/K2-Horizon) and the [vLLM recipe](https://recipes.vllm.ai/IFM).
4.  **Parsers.** Enable the `k2_horizon` reasoning parser for chat, and add the `k2_horizon` tool-call parser for agent use. Leave both off for plain completion-style generation.

## Citation

```
@misc{k2horizon2026,
  title  = {Introducing K2 Horizon: Frontier Performance, Radically Open},
  author = {{IFM Team}},
  year   = {2026},
  url    = {https://ifm.ai/blog/k2/},
}
```

## Want more deterministic results?

[Try Interfaze](https://interfaze.ai/dashboard)[Read the Docs](https://interfaze.ai/docs)
