Metallama familyActive

Llama 3.1 8B

Compact Llama instruction model for fast chat and local deployment

JSONEmbed chartCompare
Context
131,072
Max output
131,072
Input / 1M
$0.05
Output / 1M
$0.08
Providers
2
01Profile

Profile

Released
2024-07-23
Last updated
2024-07-23
Knowledge cutoff
2023-12
Weights
Open
02Capabilities

Capabilities

Input
text
Output
text
Reasoning
No
Tool calling
Yes
Structured output
No
Attachments
No
Cache read
Cache write
03Pricing

Pricing across 2 providers

Per 1M tokens, USD. Sorted by input price. First-party row highlighted.

ProviderProvider model idInputOutputCache readContext
heliconellama-3.1-8b-instant$0.05$0.08Not listed131,072
groqllama-3.1-8b-instant$0.05$0.08Not listed131,072
04History

Change history

5+ recent events since 2024-07-23

2026-01-04ContextgroqMax output8,192→ 131,072
2025-12-07Removedcloudflare-ai-gatewayofferingremoved→ new
2025-12-06Addedcloudflare-ai-gatewayoffering→ new
2025-10-23Addedheliconeoffering→ new
2025-06-12Addedgroqoffering→ new
as of 2026-09-16
About ·Corrections ·Contact ·Privacy ·EN / 中文 / 繁體
HomeModelsLabsProvidersToolsRankingsChangesNews