Skip to main content
Talk to Sales

Benchmarks

Performance, side by side.

Independent BenchLM leaderboard scores for frontier and open-weight models — so you can pick the right model for the job, not just the biggest name.

13
models compared
84.4
top score (BenchLM)
1.05M
largest context window

Frontier models

Official OpenAI & Anthropic models, provided by Tara Cloud as a reseller.

The same official APIs, on the same terms — one API key, unified JPY billing, and a single invoice through Tara Cloud.

  1. 01

    Claude Fable 5.1

    Anthropic

    84.4

    Context

    1M

  2. 02

    GPT-6 Astra

    OpenAI

    84.1

    Context

    1.05M

  3. 03

    Claude Opus 5

    Anthropic

    81.6

    Context

  4. 04

    GPT-5.6 Sol

    OpenAI

    80.5

    Context

    1.05M

  5. 05

    GPT-5.4

    OpenAI

    70.8

    Context

    1.05M

  6. 06

    Claude Sonnet 5

    Anthropic

    69.8

    Context

    1M

Official APIs resold by Tara Cloud, with unified JPY billing.

Performance per need

Match the model to the workload.

85% of the leader's score

Qwen3.8 Max delivers 85% of the leader's benchmark performance with a 1M context window — a strong fit for high-volume production workloads.

78% of the leader's score

GLM-5.3-Flash reaches 78% of the leader's score with a compact footprint — built for throughput-sensitive pipelines.

Open-weight models

Served on Tara Cloud GPUs, in Japan.

Open-weight models hosted on Tara Cloud's own GPU infrastructure in Japan — every prompt and completion stays in-country.

  1. 01

    Qwen3.8 Max

    Alibaba

    71.6

    Context

    1M

  2. 02

    GLM-5.3

    Z.AI

    68.3

    Context

    1M

  3. 03

    GLM-5.3-Flash

    Z.AI

    66.0

    Context

    1M

  4. 04

    Kimi K2.6

    Moonshot AI

    65.4

    Context

    256K

  5. 05

    MiniMax M3

    MiniMax

    61.5

    Context

    1M

  6. 06

    DeepSeek V3.2

    DeepSeek

    56.9

    Context

    128K

  7. 07

    GPT-OSS 120B

    OpenAI

    45.7

    Context

    128K

All data stays in Japan.

Benchmark data: BenchLM (benchlm.ai), snapshot of September 8, 2026. Scores are indicative and updated periodically. benchlm.ai

Pick your models. We'll handle the rest.

Tell us about your workloads and we'll size the right mix of open-weight and frontier models — usually with a proposal within two business days.