sfm baseline filtered 8k 9k 10k simpleavg merge is a language model from Yuhengtu Bytedance. It has been superseded and no longer holds a position on ModelCap's current board.
Figures as of 23 Sept 2026, 20:40 UTC · download the public dataset · methodology
Superseded by sfm baseline filtered 8k 9k 10k 11k 12k simpleavg merge. It no longer holds a current-board position.
OpenRouter endpoint status could not be refreshed. This model is not treated as currently served until a successful observation arrives.
Providers
Uptime measured over the last 30 minutesCurrent provider status is unavailable.
Other versions (10)
- sfm baseline filtered 8k 9k 10k 11k 12k simpleavg mergeEquivalent route22 days ago
- sfm baseline filtered 7k 8k 9k 10k 11k simpleavg mergeSuperseded22 days ago
- sfm baseline filtered 0k 1k 2k 3k 4k simpleavg mergeSuperseded22 days ago
- sfm baseline filtered 10k 11k 12k simpleavg mergeSuperseded25 days ago
- sfm baseline filtered 9k 10k 11k simpleavg mergeSuperseded25 days ago
- sfm baseline filtered 7k 8k 9k simpleavg mergeSuperseded25 days ago
- sfm baseline filtered 6k 7k 8k simpleavg mergeSuperseded25 days ago
- sfm baseline filtered 5k 6k 7k simpleavg mergeSuperseded25 days ago
- sfm baseline filtered 4k 5k 6k simpleavg mergeSuperseded25 days ago
- sfm baseline filtered 3k 4k 5k simpleavg mergeSuperseded25 days ago
ModelCap Index · All-versions archive
Methodology- Score range
- 27.4–87.0
Technical details
- Public rank basis
- Modeled · prior (architecture)
- Evidence label
- Modeled · prior · global-corpus-prior over 223 held-out anchors (100%)
- Index support
- 4.5% · weak
- Identity binding
- Exact catalogue product · yuhengtu-bytedance/yuhengtu-bytedance/sfm_baseline_filtered-8k_9k_10k_simpleavg_merge · aggregates disclosed benchmark configurations · endpoint configuration not claimed · artifact metadata yuhengtu-bytedance/sfm_baseline_filtered-8k_9k_10k_simpleavg_merge@c61a5bf3cd347c9c51890945f6a4c956451fc718 (not the evaluation revision)
- Index observations used
- 0
- Qualified Index evidence point
- 57.2
No general-family board lists this model yet. Its point estimate is the precision-weighted mean of every channel that spoke for it; each channel’s share is the weight it carried. Independent measurements replace this placement as boards list the model.
- Fused estimate
- 57.2 ± 23.3
- Independent share
- 0%
- Shrink toward corpus
- 0.0
- Optimism haircut
- None applied
- Method
- precision-weighted-v1
BES normalizes admitted public benchmark evidence for measured Index rows. Its legacy score and observables prior do not define the public language rank.
- Observed capability
- —
- Evidence status
- No admitted measured evidence
- General preference
- —
- Coding
- —
- Agents & tools
- —
- Reasoning
- —
- Evidence breadth
- 0%
- Evidence coverage
- 0%
- Usage (OpenRouter popularity)
- —
- Liquidity (providers × uptime)
- —
- Open reach (HF downloads)
- —
- Surface (context / tools / modalities)
- 0.0
- Freshness
- 90.6
- · Catalogue surface
- · First seen 2026-08-29
These adoption and deployment observations remain context only. They do not change this model's ModelCap Index score or rank.
The bounded 0–100 capability score blends each source’s competitive placement with its published achievement, then combines capability families. Evidence breadth adds a modest uncertainty adjustment; price and popularity are not part of benchmark-led rank.
Experimental capability estimate
Estimate policy- Confidence
- 74%
- Training support
- 232 models · 133 lineages
- Out-of-domain check
- in domain
- Feature coverage
- 71%
- Lineage-held-out error
- 9.3 MAE
Experimental estimate. This layer predicts from admitted evidence and safe metadata only. It never enters ModelCap Score or rank.
Deployment readiness
Metadata index- Artifact reproducibility
- 85.0
- Access & legal clarity
- 32.5
- Deployability
- 8.0
- Evaluation provenance
- 0.0
Not yet published: the base model it derives from, a declared licence, a link to the licence terms, a current serving provider, measured provider uptime, published provider pricing, declared context and output limits, admitted benchmark results, results across several capability areas, results from more than one source, a confident match to its benchmark entries, pinned evaluation dataset versions.
A metadata completeness and deployability index—not a safety certification, quality grade, or production-readiness claim.
Market activity
Venue coverage- Market Gravity
- 12.0
- OpenRouter weekly popularity
- Not listed
Market Gravity uses OpenRouter’s full-catalogue popularity order. Sparse Vercel top-ten observations appear only when published and do not affect the score.
Arena preference
This model is visible because it is available in the live catalogue, but it has no trustworthy community result yet. ModelCap does not infer a quality score from price, features, or market activity.
Market Gravity breakdown
- Usage55%
- 0.0
- Liquidity25%
- 0.0
- Open reach15%
- 50.0
- Freshness5%
- 90.6
OpenRouter weekly popularity across the full model catalogue.
Independent providers versus the model's open or closed cohort, weighted by uptime.
Hugging Face 30-day downloads for open models; neutral for closed models.
Time since first public availability, on a six-month half-life.
Specification
- Model ID
yuhengtu-bytedance/sfm_baseline_filtered-8k_9k_10k_simpleavg_merge- Inputs
- Text
- Outputs
- Text
- Cached input tokens
- Not offered
- First seen on OpenRouter
- 29 Aug 2026
Capabilities
- Not supported: Reasoning
- Not supported: Tool use
- Not supported: Structured output
- Not supported: Response format
- Not supported: Moderated
Weights & access
License unverified. A weight repository exists, but ModelCap could not verify its license classification from the current source metadata.
- Access
- License unverified
- Pinned revision
- c61a5bf3cd34
- Parameters
- 6.9B
- Architecture
- GPTNeoXForCausalLM
- Model type
- gpt_neox
- Hugging Face downloads (30d)
- 496
- Downloads (all time)
- 496
- Likes
- 0
- Repository updated
- 1 Sept 2026, 02:34 UTC