MetaactiveOpen Source

Llama 3.1 405B

llama-3.1-405b

Largest open-source LLM with 405B parameters.

Context Window

131.1K

tokens

Max Output

8.2K

tokens

Input Price

—

per 1M tokens

Output Price

—

per 1M tokens

Details

Familyllama-3

Parameters405B

Training Cutoff2024-12-01

ReleasedJuly 23, 2024

Aliasesmeta-llama/Llama-3.1-405B-Instruct

Capabilities

FunctionsStreamingJSON ModeCodeTool Use

Documentation

Evaluation Scores(5 benchmarks)

HumanEvalFunction-level Python code generation

89%

MATH-500Competition-style math

73.8%

MMLU-ProHarder successor to MMLU

73.3%

GPQA DiamondPhD-level science questions

50.7%

SWE-bench VerifiedResolving real GitHub issues

33.5%

Quick Access

curl pikaainews.com/api/models/meta-llama-3-1-405b

npx pika-models info meta-llama-3-1-405b

Get API Access

Official

AWS Bedrock

Amazon Bedrock. Official partnerships with Anthropic, Meta, Mistral, Cohere.

Third-Party Providers & Aggregators

Cerebras

Wafer-scale inference. 1000+ tokens/sec for select models.

DeepInfra

Lowest per-token rates for open-source models.

Fireworks AI

Fastest inference engine. Multimodal support, HIPAA/SOC2.

Groq

Ultra-fast LPU inference. Best latency for real-time apps.

OpenRouter

500+ models, one API key. Pay-per-token, no minimums.

SiliconFlow

China-optimized inference. Strong Qwen/DeepSeek support.

Together AI

Fast open-source model inference. Sub-100ms latency.