Skip to content
ModelCap

Head-to-head comparison

Trinity Large Thinking vs Qwen3.5-9B

Trinity Large Thinking (Arcee AI) and Qwen3.5-9B (Qwen) compared on the ModelCap Index, API price, context window, provider availability, weight access and every public benchmark board they share. Figures are the same ones shown on the live rankings; nothing here is a hidden score.

Snapshot as of 9 September 2026

As of 9 September 2026, Qwen3.5-9B holds the stronger ModelCap Index position (#97 vs #100); Qwen3.5-9B is 5.3× cheaper per output token ($0.15/1M vs $0.80/1M); and Qwen3.5-9B ships open weights.

Which should you choose?

Choose Trinity Large Thinking if…

  • AA τ²-Bench Telecom is your yardstick — 90.1% against 86.8%.
  • AA Humanity's Last Exam is your yardstick — 15.8% against 14.9%.
  • you want the more recently listed model — Trinity Large Thinking was listed 1 April 2026, Qwen3.5-9B 10 March 2026.

Choose Qwen3.5-9B if…

  • you want the stronger overall ModelCap Index position — #97 against #100 (41.1 vs 40.2 points).
  • API cost matters — $0.15/1M output tokens against $0.80/1M, about 5.3× cheaper.
  • your workload is prompt-heavy — input tokens cost $0.10/1M against $0.25/1M.
  • you want to self-host: Qwen3.5-9B ships open weights while Trinity Large Thinking is restricted license.
  • you want provider choice — 6 listed API providers against 1.
  • AA Terminal-Bench 2.1 is your yardstick — 29.2% against 20.6%.

Trinity Large Thinking vs Qwen3.5-9B: specs, pricing and context

Specification comparison of Trinity Large Thinking and Qwen3.5-9B
FieldTrinity Large ThinkingArcee AIQwen3.5-9BQwen
ModelCap Index position#100#97
Index score40.241.1
EvidenceMeasuredEstimated
Input price / 1M tokens$0.25$0.10
Output price / 1M tokens$0.80$0.15
Context window262K tokens262K tokens
Max output tokens80K236K
API providers listed16
Weight accessRestricted licenseOpen weights
Input modalitiesTextText, Image, Video
First listed1 April 202610 March 2026
PublisherArcee AIQwen

Evidence: 3 public benchmark observations across 3 boards · publisher-corpus-prior over 182 held-out anchors (30%); launch card against 3 resolved peers on 10 rows (70%); exceeds every named peer on 4 of 10 rows; optimism haircut 0.6 from cross-lab-probe; shrunk 3.7 up toward the measured corpus. Prices are the lowest listed API offer per million tokens observed on the OpenRouter catalogue.

Benchmark scores: Trinity Large Thinking vs Qwen3.5-9B

Public benchmark boards where Trinity Large Thinking or Qwen3.5-9B has a published result
BoardTrinity Large ThinkingQwen3.5-9BLeads
Arena codingArena (LMArena)14141407–1422Only one result
AA Intelligence IndexArtificial Analysis10.9intelligence-index:v4.3:reasoning-unspecifiedOnly one result
AA Terminal-Bench 2.1Artificial Analysis20.6%terminal-bench:2.1:terminus-2:e2b:pass-at-1:3-repeats:source-model="Trinity Large Thinking":reasoning=true29.2%terminal-bench:2.1:terminus-2:e2b:pass-at-1:3-repeats:source-model="Qwen3.5 9B":reasoning=trueQwen3.5-9B
AA τ²-Bench TelecomArtificial Analysis90.1%tau2:telecom:dual-control:pass-at-1:3-repeats:source-model="Trinity Large Thinking":reasoning=true86.8%tau2:telecom:dual-control:pass-at-1:3-repeats:source-model="Qwen3.5 9B":reasoning=trueTrinity Large Thinking
AA Humanity's Last ExamArtificial Analysis15.8%hle:may-2025:text-only-2158:no-tools:pass-at-1:source-model="Trinity Large Thinking":reasoning=true14.9%hle:may-2025:text-only-2158:no-tools:pass-at-1:source-model="Qwen3.5 9B":reasoning=trueTrinity Large Thinking

Scores are the sources' own published figures for each model's best evaluated configuration; ModelCap never re-runs a benchmark.

Want a different pairing? Open the interactive comparison tool to swap either model for any current ranked language model.

Trinity Large Thinking vs Qwen3.5-9B: common questions

Is Trinity Large Thinking better than Qwen3.5-9B?

Qwen3.5-9B ranks higher on the ModelCap Index as of 9 September 2026: #97 against #100. That is a capability ranking built from public benchmark evidence with published uncertainty; whether it is "better" for you also depends on price, context and where you can run it. Their published uncertainty intervals overlap, so the rank difference alone does not establish a reliable capability advantage for your workload.

Is Trinity Large Thinking cheaper than Qwen3.5-9B?

Qwen3.5-9B is cheaper on output tokens: $0.15/1M against $0.80/1M. Input tokens are $0.25/1M for Trinity Large Thinking and $0.10/1M for Qwen3.5-9B. Prices are the lowest listed API offer ModelCap observed, in USD per million tokens. For 1,000 requests with 2,000 input and 500 output tokens each (2M input + 0.5M output), the listed-rate estimate is $0.90 for Trinity Large Thinking versus $0.27 for Qwen3.5-9B. Qwen3.5-9B costs 69.4% less in this scenario. This excludes caching, batch discounts, prompt-length tiers, tool charges and retries; verify the selected endpoint before budgeting.

Which has the bigger context window, Trinity Large Thinking or Qwen3.5-9B?

Both publish a 262K-token context window.

Which is better for coding, Trinity Large Thinking or Qwen3.5-9B?

AA Terminal-Bench 2.1: Trinity Large Thinking 20.6%, Qwen3.5-9B 29.2% — Qwen3.5-9B leads.

Are Trinity Large Thinking and Qwen3.5-9B open-weight models?

Trinity Large Thinking: Restricted license. Qwen3.5-9B: Open weights. Open weights mean the checkpoint can be downloaded and self-hosted under its licence; API-only models are available solely through hosted endpoints.

Where can I run Trinity Large Thinking and Qwen3.5-9B?

ModelCap currently lists 1 API provider for Trinity Large Thinking and 6 for Qwen3.5-9B, from the OpenRouter catalogue snapshot the site serves; each model page lists the providers and their prices.

How current is this Trinity Large Thinking vs Qwen3.5-9B comparison?

Every figure comes from the sealed ModelCap dataset published 9 September 2026; the page re-renders within a minute of each data refresh, and the ModelCap Index positions are the same ones shown on the live rankings.

Explore further