HomeModels › Doubao-Seed-2.0-Pro
ByteDance · Flagship · 3.2T MoE · GPT-4 Class

Doubao-Seed-2.0-Pro API

ByteDance's flagship 3.2-trillion-parameter MoE model. GPT-4-class reasoning, 256K context, state-of-the-art Chinese understanding — via NovAI's zero-platform-fee Hong Kong gateway with an exclusive 10% distributor discount.

$0.40
Input / 1M tokens
$2.00
Output / 1M tokens
256K
Context window
0%
Platform fee
Sign Up — Get $2.00 Free Credit See All Pricing

Model overview

Doubao-Seed-2.0-Pro is the flagship closed-source model in ByteDance's Doubao Seed family, released in April 2026. It is a sparse Mixture-of-Experts (MoE) transformer with 3.2 trillion total parameters and approximately 280B active parameters per forward pass, trained on a 15T-token multilingual corpus with heavy emphasis on Chinese-language and bilingual reasoning data.

It occupies the top-tier of ByteDance's lineup above Doubao-Seed-2.0-Lite (for high-throughput inference) and Doubao-Seed-2.0-Code (code-specialized). Where Lite is tuned for cost-per-token and Code for HumanEval-class tasks, Pro is the default choice when you want broad-capability reasoning on complex, open-ended, Chinese-first workloads.

On NovAI, Doubao-Seed-2.0-Pro is served through a direct distributor relationship with ByteDance Volcano Engine. We pass through the entire upstream capability surface — 256K context, streaming, function calling, JSON mode, vision inputs — with no schema changes required. Requests are routed via our Hong Kong SAR point-of-presence for sub-500ms first-token latency to Asian users and sub-second latency globally.

Why use Doubao-Seed-2.0-Pro on NovAI?

  • GPT-4-class general reasoning — matches GPT-4o on MMLU (87.2 vs 88.7) and beats it on C-Eval (89.4 vs 82.1) and CMMLU (87.9 vs 83.2), making it the strongest choice for Chinese-first products.
  • 256K context window — enough for an entire 500-page report, a full monorepo, or 40 rounds of tool-use traces without manual chunking.
  • Native bilingual excellence — Chinese and English train on roughly equal token budgets; no quality degradation when switching languages mid-conversation.
  • Strong coding capability — HumanEval 88.4, MBPP 83.1, LiveCodeBench v5 62.8, competitive with GPT 3.5 Sonnet at a fraction of the output-token cost.
  • Exclusive 10% distributor discount — NovAI's direct Volcano Engine partnership lets us price Pro at $0.40/$2.00 per 1M tokens vs the official $0.44/$2.22
  • Zero platform fee — you pay only the token price. No monthly seat fee, no minimum commit, no setup cost.
  • OpenAI-compatible endpoint — change base_url and model only; no client-library migration. Works with OpenAI Python, Node.js, Go, LangChain, LlamaIndex, Vercel AI SDK, Cursor, Cline, and any tool that speaks Chat Completions.
  • Zero training opt-out — prompts and completions are never used to train models. Only aggregate billing metadata is retained.
  • Hong Kong low-latency routing — P50 time-to-first-token is 280ms from mainland China / HK / Taiwan / SEA, 450ms from Tokyo, 820ms from US-West, 1.1s from Europe.
  • USDT (TRC20) billing — prepaid credits in USD, no Chinese RMB bank account required, no KYC beyond email.

Technical specifications

PropertyValue
Model IDdoubao-seed-2.0-pro
ArchitectureSparse Mixture-of-Experts (MoE) transformer
Total parameters~3.2T
Active parameters / token~280B
Context window (input)262,144 tokens (256K)
Max output tokens16,384
TokenizerByteDance BPE (shared Chinese/English/code vocabulary, ~151K tokens)
Training cutoffDecember 2025
Release dateApril 2026
ModalitiesText in, text out. Vision (image input) available.
Streaming✓ Server-Sent Events
Function / tool calling✓ Parallel tool calls, OpenAI-compatible
JSON mode✓ Structured output via response_format
Temperature range0.0 – 2.0 (default 1.0)
Top-p range0.0 – 1.0 (default 0.7)
Rate limit (default)60 RPM · 150K TPM
Rate limit (Scale tier)600 RPM · 2M TPM (unlock at $50 balance)
SLA99.9% monthly uptime

Pricing detail

Per-token rate card

TierInput (per 1M tokens)Output (per 1M tokens)
Official Volcano Engine$0.44$2.22
NovAI (with 10% discount)$0.40$2.00
You save10%10%

Cost comparison vs competitor flagships (per 1M output tokens)

ModelOutput pricevs Doubao-Pro
GPT-4o (OpenAI)$15.007.5× more expensive
GPT 3.5 Sonnet (Anthropic)$15.007.5× more expensive
Gemini 1.5 Pro (Google)$10.505.25× more expensive
Qwen3.6-Max (Alibaba)$6.243.12× more expensive
Doubao-Seed-2.0-Pro (NovAI)$2.00baseline
DeepSeek-V4-Pro (NovAI)$0.400.2× (cheaper, weaker on Chinese)

Real-world cost estimates

  • 1M input + 200K output / day~$0.80/day, $24/month. Typical for a mid-size RAG chatbot with 500–1,000 DAU.
  • 5M input + 1M output / day~$4.00/day, $120/month. Heavy agent / code assistant usage.
  • 20M input + 4M output / day~$16/day, $480/month. Enterprise-scale deployment with long-context document pipelines.

All prices shown in USD. Billing is prepaid — credits never expire. No hidden platform fee, no minimum top-up. First-time users get $2.00 free credit on signup, enough for roughly 1.25M input tokens or 125K output tokens.

Best use cases

Chinese RAG

Enterprise Chinese-language RAG

Doubao's native Chinese excellence plus the 256K context window means you can stuff entire policy documents, HR handbooks, or research reports into a single call — no vectorstore chunking dance. Typical retrieval quality is 8–12 points higher than embedding-based hybrid search with GPT-4o.

Code assistant

Full-repo code review & refactor

256K is enough for a 50-file TypeScript backend. Point it at a whole pull request, ask for architectural critique, and it will trace state across files reliably. Pairs well with Cursor and Cline out of the box.

Agent workflow

Multi-turn agent orchestration

With parallel tool calls and JSON mode, Doubao-Pro runs complex 20+ step planner-executor loops with low error accumulation. Benchmarks on τ-bench and WebArena place it within 3 points of GPT 3.5 Sonnet at a quarter the output cost.

Long-document QA

Research / legal / finance analysis

256K handles most 10-K filings, full legal contracts, or 300-page research monographs in one shot. Needle-in-a-haystack retrieval accuracy stays above 94% up to the 240K mark.

Content creation

Marketing, copywriting, translation

Native-feeling Chinese copywriting, culturally adapted bilingual translation, and tone-matched brand content. Outperforms GPT-4o on human-preference evaluations for xiaohongshu / WeChat-style content.

Data analysis

Analytical reasoning over tables

Strong at reading pasted CSV / TSV / markdown tables and reasoning over them. Use JSON mode with a typed schema for reliable structured-extraction pipelines.

Quick start

cURL

curl https://aiapi-pro.com/v1/chat/completions \
  -H "Authorization: Bearer $NOVAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "doubao-seed-2.0-pro",
    "messages": [
      {"role":"system","content":"You are a helpful assistant."},
      {"role":"user","content":"Explain MoE architectures in 2 paragraphs."}
    ],
    "temperature": 0.7,
    "max_tokens": 800
  }'

Python (OpenAI SDK)

from openai import OpenAI

client = OpenAI(
    base_url="https://aiapi-pro.com/v1",
    api_key="YOUR_NOVAI_API_KEY",
)

resp = client.chat.completions.create(
    model="doubao-seed-2.0-pro",
    messages=[
        {"role": "system", "content": "You are a senior Python engineer."},
        {"role": "user",   "content": "Write a FastAPI health-check endpoint."},
    ],
    temperature=0.3,
    max_tokens=500,
)
print(resp.choices[0].message.content)
print("Usage:", resp.usage)

Python — streaming

stream = client.chat.completions.create(
    model="doubao-seed-2.0-pro",
    messages=[{"role":"user","content":"Write a haiku about Hong Kong."}],
    stream=True,
)
for chunk in stream:
    delta = chunk.choices[0].delta.content
    if delta:
        print(delta, end="", flush=True)

Python — function calling

tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Get current weather for a city",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
        },
    },
}]

resp = client.chat.completions.create(
    model="doubao-seed-2.0-pro",
    messages=[{"role":"user","content":"What's the weather in Hong Kong?"}],
    tools=tools,
    tool_choice="auto",
)
tool_call = resp.choices[0].message.tool_calls[0]
print(tool_call.function.name, tool_call.function.arguments)

Node.js / TypeScript

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://aiapi-pro.com/v1",
  apiKey: process.env.NOVAI_API_KEY,
});

const resp = await client.chat.completions.create({
  model: "doubao-seed-2.0-pro",
  messages: [{ role: "user", content: "Top 3 SEO tips for a SaaS landing page." }],
  temperature: 0.6,
});

console.log(resp.choices[0].message.content);

JSON mode (structured output)

resp = client.chat.completions.create(
    model="doubao-seed-2.0-pro",
    messages=[
        {"role":"system","content":"Extract a product object as JSON."},
        {"role":"user",  "content":"The new iPhone Mini is $799, releases Sep 2026, colors: red, blue."},
    ],
    response_format={"type":"json_object"},
)
# Returns: {"name":"iPhone Mini","price_usd":799,"release":"2026-09","colors":["red","blue"]}

Full documentation: aiapi-pro.com/#docs · docs.aiapi-pro.com

Performance & latency

MetricP50P95P99
TTFT (Hong Kong / CN / SEA)280ms420ms690ms
TTFT (Tokyo / Seoul)450ms680ms1.0s
TTFT (US West)820ms1.2s1.9s
TTFT (Europe)1.10s1.6s2.3s
Throughput (tokens/sec)523828
Long-context (>128K) TTFT2.1s3.4s5.6s

Measurements aggregated from 30 days of production traffic across >4,000 distinct customers. SLA commitment: 99.9% monthly uptime. If Volcano Engine upstream degrades, NovAI provides automatic fallback to Doubao-Seed-2.0-Lite with proactive status-page notification.

Doubao Seed 2.0 family — which one should I pick?

ModelBest forContextInput / 1MOutput / 1M
Doubao-Seed-2.0-ProFlagship reasoning, Chinese-first RAG, complex agents256K$0.40$2.00
Doubao-Seed-2.0-LiteHigh-throughput classification, summarization, cheap RAG128K$0.075$0.45
Doubao-Seed-2.0-CodeCode generation / completion / review at SOTA HumanEval128K$0.15$1.00

Quick rule of thumb:

  • Default to Pro when the workload mixes reasoning, long context, Chinese, and tool use.
  • Drop to Lite when volume × task simplicity pushes cost-per-output to the top of your concerns.
  • Use Code for IDE-plugin use cases (Cursor/Cline/Continue) where fill-in-the-middle and function signatures matter.

Frequently asked questions

Q: What is Doubao-Seed-2.0-Pro?

ByteDance's flagship Mixture-of-Experts large language model released in 2026. 3.2T total parameters, ~280B active per token, 256K context, GPT-4-class capability with native Chinese excellence.

Q: How much does it cost?

On NovAI: $0.40/1M input, $2.00/1M output — a 10% discount vs official Volcano Engine pricing, zero platform fee, credits never expire.

Q: Is the API OpenAI-compatible?

Yes. Endpoint is https://aiapi-pro.com/v1/chat/completions. Drop-in replace base_url and model. Supports streaming, tool calls, JSON mode, parallel function calling.

Q: What is the context window?

256K input tokens (262,144). Max output per response: 16,384 tokens. Suitable for full-report QA, monorepo analysis, or 40+ round agent workflows.

Q: Does it support streaming?

Yes, set stream: true. First-token latency is 280–420ms from Hong Kong / China / SEA.

Q: Does it support tool / function calling?

Yes — OpenAI-compatible tool calls with parallel invocation and structured JSON response.

Q: How does Doubao-Pro compare to GPT-4o and GPT 3.5 Sonnet?

Matches or beats GPT-4o on Chinese benchmarks. Trails GPT 3.5 Sonnet by ~2 points on HumanEval. ~75% cheaper per output token than both.

Q: Will my prompts be used to train future models?

No. NovAI does not retain prompts or completions for training. Only billing metadata (timestamp, token count) is logged.

Q: What latency should I expect?

P50 TTFT 280ms from Hong Kong, 450ms from Tokyo, 820ms from US West. Global throughput averages 42–58 tokens/sec.

Q: How do I get started?

Sign up at aiapi-pro.com, receive $2.00 free credit, grab your API key, point your OpenAI SDK at our base URL. Takes under 90 seconds.

Ready to try Doubao-Seed-2.0-Pro?

Zero platform fee. Credits never expire. OpenAI-compatible API — no code changes needed. $2.00 free credit on signup.

Sign Up Free Compare All Models

Related: Doubao-Seed-2.0-Lite · Doubao-Seed-2.0-Code · DeepSeek-V4-Pro · Qwen3.6-Max · GLM-5.1