Research
SGAnalog: An End-to-End Circuit Benchmark from Open-Source Silicon Tapeouts
Overview Research area: AI for electronic design automation (EDA), specifically benchmarking large language and multimodal models on analog integrated circuit design tasks. Technical level: Advanced —

- arXiv
- 2610.03934
- Published
- 2026-10-02
- Authors
- Yueting Li, Weihang Ding
AI summary
Overview
Research area: AI for electronic design automation (EDA), specifically benchmarking large language and multimodal models on analog integrated circuit design tasks.
Technical level: Advanced — the paper assumes familiarity with SPICE netlists, schematic capture, device sizing (W/L), process design kits (PDKs), and graph isomorphism.
Scope: This paper introduces SGAnalog, a commit-pinned benchmark of 273 topologically distinct analog circuit designs harvested from open-source Tiny Tapeout submissions, used to evaluate schematic-to-netlist transcription and simulator-scored device sizing across seven models.
What This Paper Is About
Existing analog circuit benchmarks make it hard to tell whether a model has learned transferable circuit skills or has simply memorized familiar textbook examples, and they rarely provide the testbenches and PDK context needed to check whether a model's output actually works. SGAnalog addresses both gaps by building a benchmark from human-designed, open-source circuits tied to Tiny Tapeout manufacturing shuttles, where every source is retrieved at the exact repository revision recorded for its shuttle submission. The goal is to evaluate two end-to-end tasks — reading a schematic into a netlist, and choosing device sizes that preserve the author's simulated behavior — while providing machinery for probing possible training-data exposure.
Key Contributions
-
A license-verified, commit-pinned dataset built offline. The benchmark contains 273 unique device–net graphs derived from 425 top-level designs, harvested from 27 Tiny Tapeout shuttles spanning 2023–2026, using a deterministic containerized pipeline with graph-hash grouping to identify redundant designs.
-
Two scored end-to-end tasks. Task 1 is graph-scored schematic transcription, where each schematic image and its reference netlist are exported from the same source file, so the image corresponds exactly to its reference. Task 2 is simulator-scored device sizing, where the model receives the human topology with sizes stripped and is graded against the human sizing under the same PDK, testbench, and simulator.
-
An exposure-aware evaluation protocol. The benchmark combines pinned commit dates with model-specific training cutoffs, blinded renders that strip author-chosen net and instance names, a BIG-bench canary in every released netlist, closed-book evaluation requests that declare no tools, and periodic refreshes from new shuttles.
-
A seven-model baseline showing the two tasks rank models differently. Transcription is led by claude-fable-5 at 56.1% exact isomorphism; sizing is led by gemini-3.1-pro-preview at 91.2 out of 100.
Main Findings
-
Transcription accuracy is low and tier-sensitive. Across seven models on a fixed set of 66 transcription tasks, the strongest model reaches 56.1% exact graph isomorphism. Six of seven models drop sharply from the small tier to the medium tier; claude-haiku-4-5 is already at floor on the small tier. Tier difficulty is not strictly monotonic: claude-opus-5 and gemini-3.1-pro-preview score higher on small circuits than on trivial ones, which are dominated by passive networks and device models.
-
Blinding labels hurts exact matches but not structure. For claude-opus-5, removing author-chosen labels lowers exact isomorphism from 51.5% to 36.4% while structural F1 is unchanged and parameter accuracy improves. The paper reads this as consistent with labels acting as visual anchors for connectivity tracing, and states it does not rule out prior exposure.
-
Failure modes differ among frontier models. claude-fable-5 leads both exact matches and structural F1 (0.853). Among the rest, claude-opus-5 is more often fully correct than gpt-5.6-sol, but its failures tend to be severe, while gpt-5.6-sol produces repairable near misses and holds the second-best structural F1 (0.831) despite a lower exact-match rate.
-
Sizing is led by a different model than transcription. On the 17 sizing tasks, gemini-3.1-pro-preview converges on all 17 proposals and reaches 91.2 out of 100 against the human reference. claude-sonnet-5 scores 85.4, gpt-5 79.0, claude-haiku-4-5 78.3, gpt-5.6-sol 78.2, claude-opus-5 74.3, and claude-fable-5 31.8.
-
Refusal, not capability, caps the newest models. claude-opus-5 returns an empty API-level refusal on 4 of the 17 sizing prompts and claude-fable-5 on 11, while both transcribe the same circuits without objection. On the tasks they answer, the two score 97.2 and 90.2. A full repeat run of claude-fable-5 refuses 12 prompts, with ten refused both times and three changing sides, so the refusal set is stable in bulk but not per prompt.
-
Convergence is too weak a bar. On one task, two proposals converge with output swing below one percent of the reference, a failure the convergence count cannot see but the simulated metric does. Every Claude or Gemini proposal that provides sizes converges, whereas both GPT generations produce several invalid device geometries.
-
Commit-date comparison shows no consistent pre-cutoff advantage. Claude-opus-5 has no post-cutoff tasks in the fixed set. Across the two usable older-cutoff comparisons, exact-match accuracy shows no consistent pre-cutoff advantage; the paper states this descriptive comparison is not a contamination estimate.
-
Static versus structured visual reading. The paper reports that a model which recovers a topology also reads its W/L annotations nearly perfectly, meaning a blended score would hide which skill is missing (W/L accuracy ranges from 0.307 to 1.000 across the seven models).
Methodology in Plain English
The authors started from the Tiny Tapeout shuttle index, which links every submission to a specific repository revision. They cloned repositories without downloading design data first, filtered candidates by file tree and license, then fetched schematics at the pinned commit — never at HEAD. Because counting every .sch file overcounts (symbols, testbenches, sub-blocks) and counting only hierarchy tops undercounts (the xschem idiom hides the interesting circuit under its testbench), they classified files by role instead: testbench, circuit, or top-level design. To measure redundancy, they computed a strict Weisfeiler–Lehman hash of each top-level device–net graph after removing net names.
Everything runs inside one pinned container image that supplies ngspice and open PDKs, so classification tables from two complete runs are byte-identical. For transcription, leaf schematics with all symbols resolved are eligible, and deterministic round-robin sampling over era and tier fixes 66 tasks. For sizing, 412 author-testbench runs were audited; a run qualifies only if it completes within 600 seconds, contains real device cards, and resolves the schematic hierarchy, leaving 190 runs, from which device-under-test mapping and a primitive-only hierarchy requirement yield 17 tasks. Scoring for transcription uses explicit graph isomorphism on a device–net bipartite multigraph, treating MOSFET drain and source as interchangeable and ignoring net names, plus a structural F1 and W/L accuracy. Sizing is scored by simulating both the model proposal and the human reference in the author's unmodified testbench and taking min(r, 1/r) per metric.
Why This Matters
The work pushes analog circuit evaluation from "did the model produce plausible text" toward "does the graph match, and does the circuit simulate correctly under the original conditions." By grounding each task in a dated, license-verified, human-authored design, it gives researchers a way to separate capability from recall in a domain where canonical textbook examples are almost certainly in pretraining data.
Real-world applications:
- Analog design automation. A model that can reliably transcribe schematics and propose workable device sizes could assist engineers in porting or documenting existing blocks.
- Verification and layout-versus-schematic checking. The scoring method used here — device–net graph isomorphism — is the same comparison at the core of an LVS check, making netlist extraction a natural fit for automated verification flows.
- Benchmark auditing and contamination analysis. The commit-date, blinded-render, and canary machinery offers a template for other benchmarks built on public repositories.
- Open-source silicon education and reuse. The released simulation browser lets anyone replay verified circuit–testbench pairs, which is useful for teaching and for building on existing Tiny Tapeout designs.
Industry relevance comes from the fact that the benchmark grades against a fixed PDK and the original author's testbench, rather than against stylistic preference, and from its finding that safety policy — not capability — sets the score on the newest models, which matters for anyone deploying these systems in a design flow.
Future Directions
- Standardized per-class measurement templates. Author testbenches expose heterogeneous measurements, so the authors plan templates starting with gain, gain-bandwidth product, and phase margin for the amplifier class, which would widen sizing coverage beyond the leaf-subcircuit set.
- Per-circuit memorization probes. The paper lists these as an archival milestone that would join the delivered blinded-render comparison.
- Topology generation as a task. The authors deliberately do not claim it in this version because two components are unsolved: a uniform representation for the design brief, and automatic pin-matching between a generated circuit and the author testbench.
- Broader process coverage. Most entries target Sky130; coverage of GF180 and IHP SG13G2, including SiGe, is planned from MPW-mirror, IHP, and ISHI-KAI sources. Reclassifying keyword-derived circuit-class labels from netlist graphs and including top-level designs without a parseable netlist are also noted.
Target Audience
Researchers and practitioners working on AI for chip design, LLM benchmarking, and analog EDA will benefit most, along with engineers interested in contamination audits for models trained on public code. The paper is written for readers comfortable with SPICE netlists, device sizing, and PDK-based simulation; readers without that background can still follow the benchmark design and the model comparisons but will need to look elsewhere for the circuit-level details.
Authors’ abstract
Existing analog integrated circuit design benchmarks make two questions hard to answer: whether a model has learned transferable circuit skills rather than recalled familiar examples, and whether its output works under defined process and test conditions. We introduce a benchmark built from human-designed, open-source circuits associated with Tiny Tapeout manufacturing shuttles. The collection contains 273 topologically distinct top-level designs. Every source is retrieved at the revision recorded for its shuttle submission and processed in a fixed containerized environment. The pipeline exports each eligible schematic image and its SPICE netlist from the same source file, giving transcription an exact structural reference. Commit dates support model-specific training-cutoff analysis, while author testbenches provide the simulation context for sizing. The benchmark evaluates schematic-to-netlist transcription and device sizing. Across seven models on a fixed set of 66 transcription tasks, the strongest model reaches 56.1% exact graph isomorphism, and six of seven models drop sharply from the small to the medium tier. For one frontier model, removing author-chosen labels reduces exact matches while preserving aggregate structural F1, suggesting that labels can aid connectivity tracing. On the 17 sizing tasks, the leading model converges on all 17 proposals and reaches 91.2 out of 100 against the human reference, while the two newest Claude models refuse 4 and 11 of the same prompts they transcribe without objection; a proposal without sizes scores zero. The two tasks produce different model rankings, exposing distinct visual and design capabilities and, in one family, a policy rather than capability limit.