{
 "cells": [
  {
   "cell_type": "markdown",
   "id": "intro",
   "metadata": {},
   "source": [
    "# Rebuild the ModelCap Index\n",
    "\n",
    "This notebook rebuilds every published ModelCap Index score and both board orders from [`/data/evidence.json`](https://modelcap.ai/data/evidence.json). It needs only the Python standard library.\n",
    "\n",
    "**What it checks**\n",
    "\n",
    "- **Measured scores.** Each board row's normalized score is weighted by its board's share of the family and by its confidence, then averaged within the family. The general family then moves by each specialist family's difference from it, capped at `specialistDeltaBound` points and weighted by that family's weight.\n",
    "- **Combined scores.** A score without a general-family board is the precision-weighted mean (weights 1/sigma squared) of the estimates in its `fusion` block. When one of those estimates is a specialist-board measurement, its mean is rebuilt from the board rows as well.\n",
    "- **Board order.** Both boards sort by score, then basis (measured, inherited, estimated), then confidence, then identity key, and number from 1.\n",
    "\n",
    "Published inputs carry one decimal, so a rebuilt score counts as matching when it lands within 0.15 points.\n",
    "\n",
    "**What it cannot check**\n",
    "\n",
    "- How a raw leaderboard result becomes a 0-100 board score. That step reads each full upstream board, which ModelCap does not redistribute. Every row links to its source.\n",
    "- How the modeled estimates in a `fusion` block are fitted: publisher corpus prior, verified lineage, launch card, corpus ladder and family succession. [The methodology](https://modelcap.ai/methodology#rank) documents them, and [launch accuracy](https://modelcap.ai/methodology#launch-accuracy) tracks how they perform.\n",
    "\n",
    "Set the `MODELCAP_EVIDENCE` environment variable to a local copy of the export to run offline. The data is licensed CC BY 4.0: credit ModelCap and link https://modelcap.ai/data."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "load",
   "metadata": {},
   "outputs": [],
   "source": [
    "import json\n",
    "import os\n",
    "import sys\n",
    "import urllib.request\n",
    "\n",
    "EVIDENCE_URL = \"https://modelcap.ai/data/evidence.json\"\n",
    "\n",
    "# A console on a legacy code page must not stop the run over a model name.\n",
    "try:\n",
    "    sys.stdout.reconfigure(errors=\"replace\")\n",
    "except (AttributeError, ValueError):\n",
    "    pass\n",
    "\n",
    "\n",
    "def load_evidence():\n",
    "    \"\"\"Read the export from MODELCAP_EVIDENCE when it is set, else fetch it.\"\"\"\n",
    "    path = os.environ.get(\"MODELCAP_EVIDENCE\")\n",
    "    if path:\n",
    "        with open(path, encoding=\"utf-8\") as handle:\n",
    "            return json.load(handle)\n",
    "    request = urllib.request.Request(EVIDENCE_URL, headers={\"User-Agent\": \"modelcap-recompute-notebook\"})\n",
    "    with urllib.request.urlopen(request, timeout=60) as response:\n",
    "        return json.load(response)\n",
    "\n",
    "\n",
    "evidence = load_evidence()\n",
    "method = evidence[\"method\"]\n",
    "models = evidence[\"models\"]\n",
    "TOLERANCE = method[\"tolerance\"]\n",
    "print(\"Snapshot:\", evidence[\"generatedAt\"])\n",
    "print(\"Rank method:\", evidence[\"methodology\"][\"indexMethodId\"])\n",
    "print(\"Evidence admission:\", evidence[\"methodology\"][\"evidenceAdmissionMethodId\"])\n",
    "print(\"Ranked models:\", len(models))"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "method-md",
   "metadata": {},
   "source": [
    "## The scoring rules\n",
    "\n",
    "These functions are the whole of the published arithmetic. Board shares, family weights and the specialist cap are read from the export's `method` block, so the notebook follows the snapshot it loads."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "method",
   "metadata": {},
   "outputs": [],
   "source": [
    "SPECIALISTS = (\"coding\", \"agent\", \"reasoning\")\n",
    "BASIS_ORDER = {\"measured\": 3, \"inherited\": 2, \"estimated\": 1}\n",
    "\n",
    "\n",
    "def clamp(value, low=0.0, high=100.0):\n",
    "    return min(high, max(low, value))\n",
    "\n",
    "\n",
    "def family_scores(observations):\n",
    "    \"\"\"Share- and confidence-weighted mean of the board scores in each family.\"\"\"\n",
    "    sums = {}\n",
    "    for row in observations:\n",
    "        weight = (row[\"share\"] or 0) * row[\"confidence\"] / 100\n",
    "        if weight <= 0:\n",
    "            continue\n",
    "        total, mass = sums.get(row[\"family\"], (0.0, 0.0))\n",
    "        sums[row[\"family\"]] = (total + row[\"sourceScore\"] * weight, mass + weight)\n",
    "    return {family: total / mass for family, (total, mass) in sums.items() if mass > 0}\n",
    "\n",
    "\n",
    "def specialist_steps(families):\n",
    "    \"\"\"Each specialist family's capped difference from general, weighted by its family weight.\"\"\"\n",
    "    bound = method[\"specialistDeltaBound\"]\n",
    "    steps = {}\n",
    "    for family in SPECIALISTS:\n",
    "        if family in families:\n",
    "            delta = clamp(families[family] - families[\"general\"], -bound, bound)\n",
    "            steps[family] = method[\"familyWeights\"][family] / 100 * delta\n",
    "    return steps\n",
    "\n",
    "\n",
    "def measured_score(observations):\n",
    "    \"\"\"The general family score, moved by each specialist family's capped difference.\"\"\"\n",
    "    families = family_scores(observations)\n",
    "    if \"general\" not in families:\n",
    "        return None\n",
    "    return clamp(families[\"general\"] + sum(specialist_steps(families).values()))\n",
    "\n",
    "\n",
    "def specialist_mean(observations):\n",
    "    \"\"\"Family-weighted mean of the specialist families a model has board rows in.\"\"\"\n",
    "    families = family_scores(observations)\n",
    "    present = [family for family in SPECIALISTS if family in families]\n",
    "    if not present:\n",
    "        return None\n",
    "    weights = [method[\"familyWeights\"][family] for family in present]\n",
    "    return sum(families[family] * weight for family, weight in zip(present, weights)) / sum(weights)\n",
    "\n",
    "\n",
    "def fused_score(fusion):\n",
    "    \"\"\"Precision-weighted mean of the estimates a combined score blends.\"\"\"\n",
    "    precision = [1 / term[\"sigma\"] ** 2 for term in fusion[\"terms\"]]\n",
    "    total = sum(weight * term[\"mean\"] for weight, term in zip(precision, fusion[\"terms\"]))\n",
    "    return clamp(total / sum(precision))\n",
    "\n",
    "\n",
    "def rebuild(model):\n",
    "    \"\"\"The score a model's published inputs support, or None when it publishes none.\"\"\"\n",
    "    if model[\"fusion\"]:\n",
    "        return fused_score(model[\"fusion\"])\n",
    "    if model[\"basis\"] == \"measured\":\n",
    "        return measured_score(model[\"observations\"])\n",
    "    return None\n",
    "\n",
    "\n",
    "def fmt(value, digits=1):\n",
    "    return \"none\" if value is None else f\"{value:.{digits}f}\""
   ]
  },
  {
   "cell_type": "markdown",
   "id": "scores-md",
   "metadata": {},
   "source": [
    "## Scores\n",
    "\n",
    "Every score is rebuilt from its own inputs. A `MISS` line names any score, family score or measured estimate that does not land within the tolerance."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "scores",
   "metadata": {},
   "outputs": [],
   "source": [
    "def within(published, rebuilt):\n",
    "    return published is not None and rebuilt is not None and abs(published - rebuilt) <= TOLERANCE\n",
    "\n",
    "\n",
    "checks = {\n",
    "    \"Measured scores from benchmark rows\": [0, 0],\n",
    "    \"Family scores from benchmark rows\": [0, 0],\n",
    "    \"Combined scores from their estimates\": [0, 0],\n",
    "    \"Measured estimates inside combined scores\": [0, 0],\n",
    "}\n",
    "unbacked = 0\n",
    "misses = []\n",
    "\n",
    "\n",
    "def check(label, published, rebuilt, miss):\n",
    "    checks[label][1] += 1\n",
    "    if within(published, rebuilt):\n",
    "        checks[label][0] += 1\n",
    "    else:\n",
    "        misses.append(miss)\n",
    "\n",
    "\n",
    "for model in models:\n",
    "    name = model[\"name\"]\n",
    "    if model[\"basis\"] == \"measured\":\n",
    "        rebuilt_families = family_scores(model[\"observations\"])\n",
    "        for family in (\"general\",) + SPECIALISTS:\n",
    "            published = model[\"families\"].get(family)\n",
    "            rebuilt = rebuilt_families.get(family)\n",
    "            if published is None and rebuilt is None:\n",
    "                continue\n",
    "            check(\"Family scores from benchmark rows\", published, rebuilt,\n",
    "                  f\"MISS {name} {family} family: published {fmt(published)}, rebuilt {fmt(rebuilt, 2)}\")\n",
    "    fusion = model[\"fusion\"]\n",
    "    if fusion:\n",
    "        check(\"Combined scores from their estimates\", model[\"score\"], fused_score(fusion),\n",
    "              f\"MISS {name}: published {model['score']:.1f}, rebuilt {fused_score(fusion):.2f}\")\n",
    "        for term in fusion[\"terms\"]:\n",
    "            if term[\"channel\"] == \"measured-specialist\":\n",
    "                mean = specialist_mean(model[\"observations\"])\n",
    "                check(\"Measured estimates inside combined scores\", term[\"mean\"], mean,\n",
    "                      f\"MISS {name} measured estimate: published {term['mean']:.1f}, rebuilt {fmt(mean, 2)}\")\n",
    "    elif model[\"basis\"] == \"measured\":\n",
    "        rebuilt = measured_score(model[\"observations\"])\n",
    "        check(\"Measured scores from benchmark rows\", model[\"score\"], rebuilt,\n",
    "              f\"MISS {name}: published {model['score']:.1f}, rebuilt {fmt(rebuilt, 2)}\")\n",
    "    else:\n",
    "        unbacked += 1\n",
    "        misses.append(f\"MISS {name}: published {model['score']:.1f}, no published inputs\")\n",
    "\n",
    "for label, (matched, checked) in checks.items():\n",
    "    print(f\"{label}: {matched} of {checked} rebuilt within {TOLERANCE} points\")\n",
    "print(f\"Scores without rebuildable inputs: {unbacked}\")\n",
    "for line in misses:\n",
    "    print(line)"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "board-md",
   "metadata": {},
   "source": [
    "## Board order\n",
    "\n",
    "Both boards are sorted from the published scores alone. A `MOVED` line names a model whose rebuilt position differs from its published one (the first ten are shown)."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "board",
   "metadata": {},
   "outputs": [],
   "source": [
    "def board(scope):\n",
    "    ranked = [model for model in models if model[\"rank\"][scope] is not None]\n",
    "    # Python compares strings by Unicode code point, which is the published tie-break.\n",
    "    ranked.sort(key=lambda model: (-model[\"score\"], -BASIS_ORDER[model[\"basis\"]], -model[\"confidence\"], model[\"identityKey\"]))\n",
    "    return ranked\n",
    "\n",
    "\n",
    "for scope, label in ((\"current\", \"Current board\"), (\"allVersions\", \"All versions\")):\n",
    "    ranked = board(scope)\n",
    "    moved = [(position, model) for position, model in enumerate(ranked, start=1) if model[\"rank\"][scope] != position]\n",
    "    print(f\"{label}: {len(ranked) - len(moved)} of {len(ranked)} positions reproduced\")\n",
    "    for position, model in moved[:10]:\n",
    "        print(f\"MOVED {model['name']}: published #{model['rank'][scope]}, rebuilt #{position}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "leaders-md",
   "metadata": {},
   "source": [
    "## Current top ten\n",
    "\n",
    "Published score beside the rebuilt one."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "leaders",
   "metadata": {},
   "outputs": [],
   "source": [
    "leaders = sorted((model for model in models if model[\"rank\"][\"current\"] is not None), key=lambda model: model[\"rank\"][\"current\"])\n",
    "print(f\"{'Rank':>4}  {'Score':>5}  {'Rebuilt':>7}  {'Basis':<9}  {'Confidence':>10}  Model\")\n",
    "for model in leaders[:10]:\n",
    "    print(f\"{model['rank']['current']:>4}  {model['score']:>5.1f}  {fmt(rebuild(model), 2):>7}  {model['basis']:<9}  {model['confidence']:>10.1f}  {model['name']}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "example-md",
   "metadata": {},
   "source": [
    "## Worked example\n",
    "\n",
    "The current #1, step by step."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "example",
   "metadata": {},
   "outputs": [],
   "source": [
    "leader = leaders[0]\n",
    "print(f\"Worked example: {leader['name']}\")\n",
    "if leader[\"fusion\"]:\n",
    "    terms = leader[\"fusion\"][\"terms\"]\n",
    "    precision = [1 / term[\"sigma\"] ** 2 for term in terms]\n",
    "    for weight, term in zip(precision, terms):\n",
    "        print(f\"  {term['channel']}: mean {term['mean']:.1f}, sigma {term['sigma']:.1f}, weight {weight / sum(precision):.3f}\")\n",
    "    print(f\"  precision-weighted mean: {fused_score(leader['fusion']):.2f}\")\n",
    "else:\n",
    "    for row in leader[\"observations\"]:\n",
    "        print(f\"  {row['family']:<9}  {row['board']}: {row['sourceScore']:.1f} x share {fmt(row['share'], 2)} x confidence {row['confidence'] / 100:.3f}\")\n",
    "    families = family_scores(leader[\"observations\"])\n",
    "    for family, value in families.items():\n",
    "        print(f\"  {family} family: {value:.2f}\")\n",
    "    for family, step in specialist_steps(families).items():\n",
    "        print(f\"  {family} adjustment: {step:+.2f}\")\n",
    "    print(f\"  rebuilt score: {measured_score(leader['observations']):.2f}\")\n",
    "print(f\"  published score: {leader['score']:.1f}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "cite-md",
   "metadata": {},
   "source": [
    "## Cite\n",
    "\n",
    "The entry names the snapshot this notebook just checked."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "cite",
   "metadata": {},
   "outputs": [],
   "source": [
    "print(evidence[\"citation\"][\"bibtex\"])"
   ]
  }
 ],
 "metadata": {
  "kernelspec": {
   "display_name": "Python 3",
   "language": "python",
   "name": "python3"
  },
  "language_info": {
   "name": "python"
  }
 },
 "nbformat": 4,
 "nbformat_minor": 5
}
