Skip to content
ModelCap

Head-to-head comparison

Trinity Large Thinking vs Step 3.7 Flash

Trinity Large Thinking (Arcee AI) and Step 3.7 Flash (StepFun) compared on the ModelCap Index, API price, context window, provider availability, weight access and every public benchmark board they share. Figures are the same ones shown on the live rankings; nothing here is a hidden score.

Snapshot as of 3 September 2026

As of 3 September 2026, Step 3.7 Flash holds the stronger ModelCap Index position (#75 vs #78); Trinity Large Thinking is 1.4× cheaper per output token ($0.80/1M vs $1.15/1M); and Step 3.7 Flash ships open weights.

Which should you choose?

Choose Trinity Large Thinking if…

  • API cost matters — $0.80/1M output tokens against $1.15/1M, about 1.4× cheaper.

Choose Step 3.7 Flash if…

  • you want the stronger overall ModelCap Index position — #75 against #78 (38.2 vs 36.4 points).
  • your workload is prompt-heavy — input tokens cost $0.20/1M against $0.25/1M.
  • you want to self-host: Step 3.7 Flash ships open weights while Trinity Large Thinking is restricted license.
  • you want provider choice — 3 listed API providers against 1.
  • AA Intelligence Index is your yardstick — 30.9 against 18.7.
  • AA Terminal-Bench 2.1 is your yardstick — 39.3% against 20.6%.
  • AA τ²-Bench Telecom is your yardstick — 98.5% against 90.1%.
  • AA Humanity's Last Exam is your yardstick — 21.4% against 15.8%.
  • you want the more recently listed model — Step 3.7 Flash was listed 28 May 2026, Trinity Large Thinking 1 April 2026.

Trinity Large Thinking vs Step 3.7 Flash: specs, pricing and context

Specification comparison of Trinity Large Thinking and Step 3.7 Flash
FieldTrinity Large ThinkingArcee AIStep 3.7 FlashStepFun
ModelCap Index position#78#75
Index score36.438.2
EvidenceMeasuredMeasured
Input price / 1M tokens$0.25$0.20
Output price / 1M tokens$0.80$1.15
Context window262K tokens262K tokens
Max output tokens80K230K
API providers listed13
Weight accessRestricted licenseOpen weights
Input modalitiesTextText, Image, Video
First listed1 April 202628 May 2026
PublisherArcee AIStepFun

Evidence: 3 public benchmark observations across 3 boards · 1 public benchmark observation across 1 board. Prices are the lowest listed API offer per million tokens observed on the OpenRouter catalogue.

Benchmark scores: Trinity Large Thinking vs Step 3.7 Flash

Public benchmark boards where Trinity Large Thinking or Step 3.7 Flash has a published result
BoardTrinity Large ThinkingStep 3.7 FlashLeads
Arena codingArena (LMArena)14141407–1422Only one result
AA Intelligence IndexArtificial Analysis18.7intelligence-index30.9intelligence-indexStep 3.7 Flash
AA Terminal-Bench 2.1Artificial Analysis20.6%terminal-bench-2.139.3%terminal-bench-2.1Step 3.7 Flash
AA τ²-Bench TelecomArtificial Analysis90.1%tau2-bench-telecom98.5%tau2-bench-telecomStep 3.7 Flash
AA Humanity's Last ExamArtificial Analysis15.8%humanitys-last-exam21.4%humanitys-last-examStep 3.7 Flash

Scores are the sources' own published figures for each model's best evaluated configuration; ModelCap never re-runs a benchmark.

Want a different pairing? Open the interactive comparison tool to swap either model for any current ranked language model.

Trinity Large Thinking vs Step 3.7 Flash: common questions

Is Trinity Large Thinking better than Step 3.7 Flash?

Step 3.7 Flash ranks higher on the ModelCap Index as of 3 September 2026: #75 against #78. That is a capability ranking built from public benchmark evidence with published uncertainty; whether it is "better" for you also depends on price, context and where you can run it.

Is Trinity Large Thinking cheaper than Step 3.7 Flash?

Trinity Large Thinking is cheaper on output tokens: $0.80/1M against $1.15/1M. Input tokens are $0.25/1M for Trinity Large Thinking and $0.20/1M for Step 3.7 Flash. Prices are the lowest listed API offer ModelCap observed, in USD per million tokens.

Which has the bigger context window, Trinity Large Thinking or Step 3.7 Flash?

Both publish a 262K-token context window.

Which is better for coding, Trinity Large Thinking or Step 3.7 Flash?

AA Terminal-Bench 2.1: Trinity Large Thinking 20.6%, Step 3.7 Flash 39.3% — Step 3.7 Flash leads.

Are Trinity Large Thinking and Step 3.7 Flash open-weight models?

Trinity Large Thinking: Restricted license. Step 3.7 Flash: Open weights. Open weights mean the checkpoint can be downloaded and self-hosted under its licence; API-only models are available solely through hosted endpoints.

Where can I run Trinity Large Thinking and Step 3.7 Flash?

ModelCap currently lists 1 API provider for Trinity Large Thinking and 3 for Step 3.7 Flash, from the OpenRouter catalogue snapshot the site serves; each model page lists the providers and their prices.

How current is this Trinity Large Thinking vs Step 3.7 Flash comparison?

Every figure comes from the sealed ModelCap dataset published 3 September 2026; the page re-renders within a minute of each data refresh, and the ModelCap Index positions are the same ones shown on the live rankings.

Explore further