n4nAI

Model catalog

Every model, one meter.

48+ models across Anthropic, OpenAI, Google, Meta, Mistral, DeepSeek and more — one endpoint, live pricing, no separate integration per provider.

prices shown per million tokens · updated on every deploy from the live catalog

48 models

GPT-5.6 Lunaby openai

openai/gpt-5.6-luna

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...

1050K context
in $0.05/M · out $0.29/M
Jul 2026
GPT-5.6 Terraby openai

openai/gpt-5.6-terra

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...

1050K context
in $0.12/M · out $0.72/M
Jul 2026
GPT-5.6 Solby openai

openai/gpt-5.6-sol

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...

1050K context
in $0.24/M · out $1.44/M
Jul 2026
Claude Sonnet 5by anthropic

anthropic/claude-sonnet-5

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...

1M context
in $2.00/M · out $10.00/M
Jun 2026
GLM 5.2by z-ai

z-ai/glm-5.2

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

1,048,576 context
in $1.39/M · out $4.40/M
Jun 2026
Kimi K2.7 Codeby moonshotai

moonshotai/kimi-k2.7-code

MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts...

262,144 context
in $0.94/M · out $4.00/M
Jun 2026
Claude Fable 5by anthropic

anthropic/claude-fable-5

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...

1M context
in $10.00/M · out $50.00/M
Jun 2026
Claude Opus 4.8by anthropic

anthropic/claude-opus-4.8

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...

1M context
in $5.00/M · out $25.00/M
May 2026

google/gemini-3.5-flash

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...

1,048,576 context
in $1.50/M · out $9.00/M
May 2026

openai/gpt-chat-latest

GPT Chat Latest points to OpenAI's stable API alias chat-latest that always resolves to the latest Instant chat model used in ChatGPT. As OpenAI rolls out new Instant model updates...

400K context
in $0.24/M · out $1.44/M
May 2026

qwen/qwen3.6-35b-a3b

Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...

262,144 context
in $0.25/M · out $1.25/M
Apr 2026

qwen/qwen3.6-27b

Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs...

262,144 context
in $0.60/M · out $3.60/M
Apr 2026
DeepSeek V4 Proby deepseek

deepseek/deepseek-v4-pro

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...

1,048,576 context
in $1.74/M · out $3.46/M
Apr 2026

deepseek/deepseek-v4-flash

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...

1,048,576 context
in $0.01/M · out $0.01/M
Apr 2026
Kimi K2.6by moonshotai

moonshotai/kimi-k2.6

Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and...

262,144 context
in $0.95/M · out $4.00/M
Apr 2026
GLM 5.1by z-ai

z-ai/glm-5.1

GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...

202,752 context
in $1.40/M · out $4.40/M
Apr 2026
Gemma 4 31Bby google

google/gemma-4-31b-it

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

262,144 context
in $0.30/M · out $1.25/M
Apr 2026

qwen/qwen3.5-35b-a3b

The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall...

262,144 context
in $0.25/M · out $1.25/M
Feb 2026

qwen/qwen3.5-27b

The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of...

262,144 context
in $0.39/M · out $3.12/M
Feb 2026

google/gemini-3.1-pro-preview

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...

1,048,576 context
in $2.00/M · out $12.00/M
Feb 2026
MiniMax M2.5by minimax

minimax/minimax-m2.5

MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1...

204,800 context
in $0.30/M · out $1.20/M
Feb 2026
Kimi K2.5by moonshotai

moonshotai/kimi-k2.5

Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm paradigm. Built on Kimi K2 with continued pretraining over approximately 15T mixed...

262,144 context
in $0.60/M · out $3.00/M
Jan 2026

meta-llama/llama-3.3-70b-instruct

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...

131,072 context
in $0.71/M · out $0.71/M
Dec 2024

meta-llama/llama-3.1-70b-instruct

Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 70B instruct-tuned version is optimized for high quality dialogue usecases. It has demonstrated strong...

131,072 context
in $0.80/M · out $0.80/M
Jul 2024

meta-llama/llama-3.1-8b-instruct

Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...

131,072 context
in $0.22/M · out $0.22/M
Jul 2024
Claude Haiku 4.5by anthropic

anthropic/claude-haiku-4-5

Anthropic fast model

200K context
in $1.00/M · out $5.00/M
Opus 5by anthropic

anthropic/claude-opus-5

1M context
in $5.00/M · out $25.00/M

google/gemini-3-pro-image

Gemini 3 Pro Image (Nano Banana Pro) — Google's highest-quality image generation model. Text-to-image at 1K/2K, served through the node-local Gemini router.

0 context
in $0.00/M · out $0.00/M
Nano Banana 2by google

google/gemini-3.1-flash-image

Gemini 3.1 Flash Image (Nano Banana 2) — fast text-to-image generation, the router's default image model.

0 context
in $0.00/M · out $0.00/M

google/gemini-3.1-flash-lite-image

Gemini 3.1 Flash Lite Image (Nano Banana 2 Lite) — cheapest and fastest text-to-image tier.

0 context
in $0.00/M · out $0.00/M

google/gemini-omni-flash

Gemini Omni Flash — fastest text-to-video tier, 10-second 720p clips.

0 context
in $0.00/M · out $0.00/M
Veo 3by google

google/veo-3.0

Veo 3 — previous-generation text-to-video with native audio, distilled for faster turnaround. 720p.

0 context
in $0.00/M · out $0.00/M
Veo 3.1by google

google/veo-3.1

Veo 3.1 — Google's text-to-video model with native audio. 720p clips, generated asynchronously.

0 context
in $0.00/M · out $0.00/M
Kimi K3by moonshotai

moonshotai/kimi-k3

MoonshotAI: Kimi K3 is the latest flagship in the Kimi family — a long-context (1M) multimodal thinking model. Served via the node-local router (kimi-code upstream).

1,048,576 context
in $5.00/M · out $25.00/M
GPT-4oby openai

openai/gpt-4o

128K context
in $0.12/M · out $0.48/M
GPT-4o miniby openai

openai/gpt-4o-mini

Fast, low-cost OpenAI model

128K context
in $0.0072/M · out $0.03/M

openai/gpt-oss-120b

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...

131,072 context
in $0.15/M · out $0.60/M

openai/gpt-oss-20b

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...

131,072 context
in $0.05/M · out $0.20/M

qwen/qwen-image-3.0

Qwen image generation v3.0, via Alibaba DashScope.

0 context
in $0.00/M · out $0.00/M

qwen/qwen-image-3.0-pro

Qwen image generation v3.0 Pro, via Alibaba DashScope.

0 context
in $0.00/M · out $0.00/M

qwen/qwen-image-edit

Qwen image editing, via Alibaba DashScope.

0 context
in $0.00/M · out $0.00/M

qwen/qwen-image-edit-max

Qwen image editing (max), via Alibaba DashScope.

0 context
in $0.00/M · out $0.00/M

qwen/qwen-image-edit-plus

Qwen image editing (plus), via Alibaba DashScope.

0 context
in $0.00/M · out $0.00/M

qwen/qwen-image-max

Qwen image generation (max quality), via Alibaba DashScope.

0 context
in $0.00/M · out $0.00/M

qwen/qwen-image-plus

Qwen image generation (plus), via Alibaba DashScope.

0 context
in $0.00/M · out $0.00/M

qwen/wan2.5-t2v-preview

Wan 2.5 text-to-video (480P), via Alibaba DashScope.

0 context
in $0.00/M · out $0.00/M

qwen/wan2.6-t2v

Wan 2.6 text-to-video (720P), via Alibaba DashScope.

0 context
in $0.00/M · out $0.00/M

qwen/wan2.7-t2v

Wan 2.7 text-to-video (720P), via Alibaba DashScope.

0 context
in $0.00/M · out $0.00/M

Open the full catalog in your console.

See live prices, context windows, and rate limits for every model.

Open console