Meta · all models ›

MLlama 4 ScoutOPEN

Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B.

1.3MContext window · tokens
$0.11Input price · per M tokens
$0.34Output price · per M tokens
MetaProvider

Prices update automatically — checked daily against provider list prices.

See model comparisons → Compare all model prices

Benchmarks & performance

Independent benchmark scores for Llama 4 Scout, measured by Artificial Analysis. Higher is better.

5.4×
More intelligence per dollar than Claude Opus 5 (Fast).
For the same budget, Llama 4 Scout delivers 5.4 times more capability. Open weights also mean you can self-host it and pay nothing per token.
Intelligence index10.3
Coding index8.2
Math index14
GPQA58.7%
MMLU-Pro75.2%
Humanity's Last Exam3.8%
Long Context Reasoning30.3%
LiveCodeBench29.9%
SciCode17%
MATH-50084.4%
AIME28.3%
AIME 202514%
IFBench39.5%
τ²-Bench15.5%
τ-Bench Banking3.3%
Terminal-Bench3.7%
Terminal-Bench Hard1.5%
⚡ Speed137.4 tokens/sec
⏱ Latency0.56s to first token
💰 Blended price$0.3 / 1M tokens
📈 Value34.3 intelligence points per $
vs. models measured heretop 82%
Scores higher than 18% of the 273 models measured by Artificial Analysis and tracked here.
Models at this level cost $0.23 per 1M tokens (median of 24) — this one costs $0.17.
Cheaper and better on this index: GLM 5.3 Flash · DeepSeek V4 Flash 0731 · DeepSeek V4 Flash 0423 and 20 more
Benchmark data by Artificial Analysis

About this model

Llama 4 Scout is an open-weight AI model by Meta. You can download and self-host it for free; the prices below are hosted-API list prices, tracked daily, for when you prefer convenience over self-hosting.

Frequently asked questions

What is Llama 4 Scout?

Llama 4 Scout is an AI language model from Meta. It is open-weight: you can download it and run it on your own hardware, for free. It scores 10.3 on the Artificial Analysis intelligence index.

Is Llama 4 Scout free?

The weights are free and open — you can self-host Llama 4 Scout and pay nothing per token. If you prefer a hosted API, list prices are $0.11 per million input tokens and $0.34 per million output tokens.

What is Llama 4 Scout good at?

Independent benchmarks from Artificial Analysis give it GPQA 58.7%, MMLU-Pro 75.2%, Humanity's Last Exam 3.8%, Long Context Reasoning 30.3%, LiveCodeBench 29.9%, SciCode 17%, MATH-500 84.4%, AIME 28.3%, AIME 2025 14%, IFBench 39.5%, τ²-Bench 15.5%, τ-Bench Banking 3.3%, Terminal-Bench 3.7%, Terminal-Bench Hard 1.5%. It is particularly used for code generation.

How fast is Llama 4 Scout?

It generates about 137.4 tokens per second, with a median 0.56s delay before the first token. Measured independently by Artificial Analysis.

Can I self-host Llama 4 Scout?

Yes. Llama 4 Scout has open weights, so you can download it and run it on your own GPU or server with tools like Ollama, vLLM or llama.cpp — with no per-token cost.

Related models

Muse Spark 1.2MetaMuse Spark 1.1MetaMuse Spark 1.2 ContributorMetaLlama 3.1 8B InstructMetaLlama 3.3 70B InstructMetaMuse Glimmer 30BMeta