AI Models
Explore available text and image models
Text Models
Image Models
Token Usage (last 30 days)
Images generated (last 30 days)
KIMI-K3
Moonshot AI flagship mixture-of-experts model with 1M-token context for coding, reasoning, and agentic workflows
MiniMax-M2.5
MiniMax model for coding, search, tool use, and office tasks
MiniMax-M3
MiniMax MoE model with 1M-token context support
Nemotron-3-Super-120B-A12B
NVIDIA 120B MoE model with 12B active parameters for reasoning, tool use, and long-context tasks
Llama 3.3 70B Instruct
Large language model from Meta
GPT-OSS 120B
Open-source alternative to GPT-4 with strong performance
Qwen3-235B-A22B-Instruct
High-performance MoE model delivering flagship-quality outputs with efficient resource usage
GLM-5.2
Advanced agentic reasoning and coding model from Z.ai
KIMI-K2.6
Moonshot AI mixture-of-experts model for coding, reasoning, and agentic workflows
DeepSeek-V4-Pro
DeepSeek flagship reasoning and coding model with long-context support
DeepSeek-R1
DeepSeek reasoning model trained with reinforcement learning for maths, code, and multi-step logic
DeepSeek-V3.2
DeepSeek general-purpose MoE model with sparse attention for long-context work
DeepSeek-V4-Flash
Low-latency DeepSeek model with 1M-token context for high-throughput workloads
Gemma-4-26B-A4B-IT
Google open MoE model with 4B active parameters for efficient general-purpose use
Gemma-4-31B-IT
Google open dense instruction-tuned model for reasoning and creative tasks
GLM-4.7
Z.ai coding and agentic model with stable multi-step execution
GLM-5
Z.ai flagship reasoning and coding model
GLM-5.1
Z.ai agentic reasoning and coding model with extended tool use
GPT-OSS 20B
Compact open-weight OpenAI model for fast, low-cost inference
KIMI-K2.5
Moonshot AI multimodal mixture-of-experts model for vision, coding, reasoning, and agentic workflows
Llama 3.1 70B Instruct
Meta instruction-tuned 70B model for general-purpose chat and reasoning
Llama 3.1 8B Instruct
Small, fast Meta instruction-tuned model for high-volume workloads
MiniMax-M2.7
MiniMax model for coding, search, tool use, and office tasks
Ministral 14B
Mistral compact instruction-tuned model for edge and low-latency use
Mistral Large 3
Mistral flagship MoE model for complex reasoning and multilingual tasks
Mistral Medium 3.5
Mistral mid-size frontier-class model balancing quality and cost
Mistral Small 4
Mistral efficient general-purpose model for everyday workloads
Nemotron-3-Ultra-550B-A55B
NVIDIA 550B MoE flagship with 55B active parameters for reasoning and long-context tasks
Qwen3.6-35B-A3B
Qwen MoE model with 3B active parameters for fast general-purpose inference
Didn't find your model?
Deploy the model of your choice as a private inference endpoint on dedicated GPUs in a few clicks.
