New Kimi K2 Thinking & Qwen3-Max are live. Browse models →

BrushLLM

Model catalog

24 models. 10 labs.One endpoint.

Every model below is callable through the same OpenAI-compatible API and the same key. Prices are indicative, in USD per 1M tokens — 3 models are free.

Showing 24 of 24 models

DeepSeek

DeepSeek-V3.2

Featured

Flagship non-thinking model. Fast, cheap, and excellent at general chat and tool use.

ChatTool CallingLong Context
Context
128K
Input / 1M
$0.28
Output / 1M
$0.42
deepseek-chatTry it

DeepSeek

DeepSeek-R1

Featured

Reasoning model with visible chain-of-thought for math, code, and hard problems.

ReasoningTool CallingLong Context
Context
128K
Input / 1M
$0.55
Output / 1M
$2.19
deepseek-reasonerTry it

Alibaba Qwen

Qwen3-Max

Featured

Alibaba’s largest flagship. Top-tier quality for demanding agentic and chat workloads.

ChatTool CallingLong Context
Context
256K
Input / 1M
$1.20
Output / 1M
$6
qwen3-maxTry it

Alibaba Qwen

Qwen3 235B A22B

Open-weights MoE flagship with hybrid thinking mode.

ChatReasoningTool CallingLong Context
Context
256K
Input / 1M
$0.20
Output / 1M
$0.80
qwen3-235b-a22bTry it

Alibaba Qwen

Qwen3 32B

Dense open-weights all-rounder with thinking and non-thinking modes.

ChatReasoningTool Calling
Context
128K
Input / 1M
$0.06
Output / 1M
$0.24
qwen3-32bTry it

Alibaba Qwen

Qwen3 Coder Plus

Code-specialized model built for agentic coding tools and repository-scale edits.

ChatTool CallingLong Context
Context
256K
Input / 1M
$0.35
Output / 1M
$1.40
qwen3-coder-plusTry it

Alibaba Qwen

Qwen VL Max

Vision-language model for image understanding, OCR, and chart reasoning.

ChatVision
Context
128K
Input / 1M
$0.28
Output / 1M
$0.83
qwen-vl-maxTry it

Alibaba Qwen

Text Embedding v4

Multilingual embedding model for search and RAG pipelines.

Embeddings
Context
8K
Input / 1M
$0.05
Output / 1M
text-embedding-v4Try it

Zhipu GLM

GLM-4.6

Featured

Z.ai flagship. Strong coding, writing, and agentic planning; 200K context.

ChatReasoningTool CallingLong Context
Context
200K
Input / 1M
$0.60
Output / 1M
$2.20
glm-4.6Try it

Zhipu GLM

GLM-4.5 Air

Lighter GLM tuned for cost-sensitive production traffic.

ChatTool Calling
Context
128K
Input / 1M
$0.20
Output / 1M
$1.10
glm-4.5-airTry it

Zhipu GLM

GLM-4.5 Flash

Free tier of the GLM family — great for prototypes and hobby projects.

ChatTool Calling
Context
128K
Input / 1M
Free
Output / 1M
Free
glm-4.5-flashTry it

Moonshot Kimi

Kimi K2 (0905)

Featured

Trillion-parameter MoE tuned for agentic tool use and long documents.

ChatTool CallingLong Context
Context
256K
Input / 1M
$0.60
Output / 1M
$2.50
kimi-k2-0905Try it

Moonshot Kimi

Kimi K2 Thinking

Reasoning variant of K2 with interleaved thinking for multi-step tasks.

ReasoningTool CallingLong Context
Context
256K
Input / 1M
$0.60
Output / 1M
$2.50
kimi-k2-thinkingTry it

ByteDance Doubao

Doubao Seed 1.6

Volcano Engine’s value flagship with controllable thinking depth.

ChatReasoningTool CallingLong Context
Context
256K
Input / 1M
$0.04
Output / 1M
$0.16
doubao-seed-1-6Try it

ByteDance Doubao

Doubao Seed 1.6 Flash

Ultra-cheap fast tier — near-free for high-volume workloads.

ChatTool Calling
Context
256K
Input / 1M
Free
Output / 1M
Free
doubao-seed-1-6-flashTry it

ByteDance Doubao

Doubao Seed 1.6 Vision

Multimodal Seed model for image understanding at Flash prices.

ChatVision
Context
128K
Input / 1M
$0.05
Output / 1M
$0.15
doubao-seed-1-6-visionTry it

MiniMax

MiniMax-M2

Efficient agentic MoE — a favorite for coding agents on a budget.

ChatTool CallingLong Context
Context
200K
Input / 1M
$0.30
Output / 1M
$1.20
MiniMax-M2Try it

Baidu ERNIE

ERNIE 4.5 Turbo

Hybrid-thinking multimodal model at aggressive prices.

ChatTool Calling
Context
128K
Input / 1M
$0.11
Output / 1M
$0.44
ernie-4.5-turboTry it

Baidu ERNIE

ERNIE 4.5 X1

Baidu’s reasoning flagship for planning and tool orchestration.

ReasoningTool Calling
Context
128K
Input / 1M
$0.28
Output / 1M
$1.10
ernie-4.5-x1Try it

Tencent Hunyuan

Hunyuan TurboS

Tencent’s hybrid fast model — switches thinking on only when needed.

ChatTool CallingLong Context
Context
256K
Input / 1M
$0.11
Output / 1M
$0.28
hunyuan-turbosTry it

Tencent Hunyuan

Hunyuan T1

Deep-thinking model with long chain-of-thought for hard reasoning.

Reasoning
Context
32K
Input / 1M
$0.28
Output / 1M
$1.10
hunyuan-t1Try it

iFlytek Spark

Spark X1

iFlytek’s reasoning model with strong multilingual understanding.

Reasoning
Context
32K
Input / 1M
$0.10
Output / 1M
$0.40
spark-x1Try it

iFlytek Spark

Spark Lite

Free lightweight tier for chat and simple tools.

ChatTool Calling
Context
32K
Input / 1M
Free
Output / 1M
Free
spark-liteTry it

StepFun

Step-3

Multimodal flagship from StepFun with interleaved thinking.

ChatReasoningVision
Context
64K
Input / 1M
$0.30
Output / 1M
$0.90
step-3Try it

Prices are indicative list prices in USD per 1M tokens and may differ from the rates configured in your console. Context windows are the maximum supported by each upstream model.