AI Models

Explore available text and image models

Text Models

29models

Image Models

0models

Token Usage (last 30 days)

...tokens

Images generated (last 30 days)

...images
Get API Key

KIMI-K3

Moonshot AI flagship mixture-of-experts model with 1M-token context for coding, reasoning, and agentic workflows

moonshotai
Pricing / M tokens
Input$3
Output$15
Cached$0.3
Documentation
Try in Playground

MiniMax-M2.5

MiniMax model for coding, search, tool use, and office tasks

MiniMaxAI
Pricing / M tokens
Input$0.3
Output$1.2
Cached$0.1
Documentation
Try in Playground

MiniMax-M3

MiniMax MoE model with 1M-token context support

MiniMaxAI
Pricing / M tokens
Input$0.3
Output$1.2
Cached$0.06
Documentation
Try in Playground

Nemotron-3-Super-120B-A12B

NVIDIA 120B MoE model with 12B active parameters for reasoning, tool use, and long-context tasks

nvidia
Pricing / M tokens
Input$0.4
Output$0.8
Cached$0.1
Documentation
Try in Playground

Llama 3.3 70B Instruct

Large language model from Meta

meta-llama
Pricing / M tokens
Input$0.7
Output$0.7
Cached$0.2
Documentation
Try in Playground

GPT-OSS 120B

Open-source alternative to GPT-4 with strong performance

openai
Pricing / M tokens
Input$0.15
Output$0.6
Cached$0.1
Documentation
Try in Playground

Qwen3-235B-A22B-Instruct

High-performance MoE model delivering flagship-quality outputs with efficient resource usage

Qwen
Pricing / M tokens
Input$0.25
Output$1
Cached$0.2
Documentation
Try in Playground

GLM-5.2

Advanced agentic reasoning and coding model from Z.ai

zai-org
Pricing / M tokens
Input$1.3
Output$4.2
Cached$0.25
Documentation
Try in Playground

KIMI-K2.6

Moonshot AI mixture-of-experts model for coding, reasoning, and agentic workflows

moonshotai
Pricing / M tokens
Input$1.2
Output$4.5
Cached$0.2
Documentation
Try in Playground

DeepSeek-V4-Pro

DeepSeek flagship reasoning and coding model with long-context support

deepseek-ai
Pricing / M tokens
Input$1.7
Output$3.4
Cached$0.15
Documentation
Try in Playground

DeepSeek-R1

DeepSeek reasoning model trained with reinforcement learning for maths, code, and multi-step logic

deepseek-ai
Pricing / M tokens
Input$1.5
Output$4.5
Cached$0.3
Documentation
Try in Playground

DeepSeek-V3.2

DeepSeek general-purpose MoE model with sparse attention for long-context work

deepseek-ai
Pricing / M tokens
Input$0.5
Output$1.5
Cached$0.15
Documentation
Try in Playground

DeepSeek-V4-Flash

Low-latency DeepSeek model with 1M-token context for high-throughput workloads

deepseek-ai
Pricing / M tokens
Input$0.2
Output$0.4
Cached$0.05
Documentation
Try in Playground

Gemma-4-26B-A4B-IT

Google open MoE model with 4B active parameters for efficient general-purpose use

google
Pricing / M tokens
Input$0.2
Output$0.5
Cached$0.1
Documentation
Try in Playground

Gemma-4-31B-IT

Google open dense instruction-tuned model for reasoning and creative tasks

google
Pricing / M tokens
Input$0.35
Output$0.8
Cached$0.2
Documentation
Try in Playground

GLM-4.7

Z.ai coding and agentic model with stable multi-step execution

zai-org
Pricing / M tokens
Input$0.7
Output$2.8
Cached$0.2
Documentation
Try in Playground

GLM-5

Z.ai flagship reasoning and coding model

zai-org
Pricing / M tokens
Input$1
Output$3
Cached$0.2
Documentation
Try in Playground

GLM-5.1

Z.ai agentic reasoning and coding model with extended tool use

zai-org
Pricing / M tokens
Input$1.4
Output$4.4
Cached$0.3
Documentation
Try in Playground

GPT-OSS 20B

Compact open-weight OpenAI model for fast, low-cost inference

openai
Pricing / M tokens
Input$0.07
Output$0.25
Cached$0.05
Documentation
Try in Playground

KIMI-K2.5

Moonshot AI multimodal mixture-of-experts model for vision, coding, reasoning, and agentic workflows

moonshotai
Pricing / M tokens
Input$0.6
Output$3
Cached$0.2
Documentation
Try in Playground

Llama 3.1 70B Instruct

Meta instruction-tuned 70B model for general-purpose chat and reasoning

meta-llama
Pricing / M tokens
Input$0.8
Output$0.8
Cached$0.16
Documentation
Try in Playground

Llama 3.1 8B Instruct

Small, fast Meta instruction-tuned model for high-volume workloads

meta-llama
Pricing / M tokens
Input$0.15
Output$0.15
Cached$0.03
Documentation
Try in Playground

MiniMax-M2.7

MiniMax model for coding, search, tool use, and office tasks

MiniMaxAI
Pricing / M tokens
Input$0.4
Output$1.5
Cached$0.15
Documentation
Try in Playground

Ministral 14B

Mistral compact instruction-tuned model for edge and low-latency use

mistralai
Pricing / M tokens
Input$0.2
Output$0.2
Cached$0.02
Documentation
Try in Playground

Mistral Large 3

Mistral flagship MoE model for complex reasoning and multilingual tasks

mistralai
Pricing / M tokens
Input$0.5
Output$1.5
Cached$0.05
Documentation
Try in Playground

Mistral Medium 3.5

Mistral mid-size frontier-class model balancing quality and cost

mistralai
Pricing / M tokens
Input$1.5
Output$7.5
Cached$0.3
Documentation
Try in Playground

Mistral Small 4

Mistral efficient general-purpose model for everyday workloads

mistralai
Pricing / M tokens
Input$0.15
Output$0.6
Cached$0.015
Documentation
Try in Playground

Nemotron-3-Ultra-550B-A55B

NVIDIA 550B MoE flagship with 55B active parameters for reasoning and long-context tasks

nvidia
Pricing / M tokens
Input$0.6
Output$3.5
Cached$0.2
Documentation
Try in Playground

Qwen3.6-35B-A3B

Qwen MoE model with 3B active parameters for fast general-purpose inference

Qwen
Pricing / M tokens
Input$0.25
Output$1.5
Cached$0.2
Documentation
Try in Playground

Didn't find your model?

Deploy the model of your choice as a private inference endpoint on dedicated GPUs in a few clicks.

Deploy your private endpoint