Skip to content
AI.info

The Pulse

GPT-5.6 Sol Completes 36 of 40 Qubit Measurements Autonomously

OpenAI says GPT-5.6 Sol, running through Codex, completed 36 of 40 target measurements on a previously uncalibrated six-qubit chip with only four researcher interventions. The MIT experiment shows where AI agents can manage routine quantum

GPT-5.6 Sol Completes 36 of 40 Qubit Measurements Autonomously

AI.info Team ·

GPT-5.6 Sol completed 36 of 40 target measurements on a previously uncalibrated superconducting-qubit chip, with researchers stepping in to improve only four. The result, reported by OpenAI and MIT’s Engineering Quantum Systems Group, marks a practical test of an AI agent controlling live laboratory equipment rather than simply analyzing an offline dataset.

OpenAI connected GPT-5.6 Sol to Codex and MIT’s existing measurement software. The agent selected parameters, operated instruments, interpreted incoming data, and chose whether to refine a measurement or proceed to the next stage. The work focused on routine calibration of a six-qubit chip, not the design or execution of a new quantum algorithm.

Forty measurements, four interventions

MIT graduate student Beatriz Yankelevich used Codex through an in-house Jupyter-based interface connected to the Engineering Quantum Systems Group’s laboratory orchestrator. The system gave the agent access to live measurement parameters, control programs, plots, raw data, logs, a measurement database and laboratory notes.

The chip contained six uncoupled qubits, including four fixed-frequency qubits and two tunable-frequency qubits. It had never been measured before. GPT-5.6 Sol first discovered the six readout resonators, identified their approximate frequencies and selected suitable pulse powers. OpenAI’s technical case study says the agent then completed a standard set of measurements on the four fixed-frequency qubits, including spectroscopy, pulse calibration, readout optimization and coherence measurements.

For one qubit, the agent measured a relaxation time of about 32 microseconds and a Ramsey coherence time of about 59 microseconds. Those values came from the calibration sequence, which also included measurements of the qubit’s transition energies, resonator response and effective temperature.

The experiment used GPT-5.6 Sol at the Ultra reasoning level. OpenAI’s technical case study, dated September 4, 2026, says the agent operated the live hardware through EQuS’s existing orchestration software, without a specialized agent framework beyond a simple in-house interface.

What the agent actually controlled

Superconducting-qubit experiments are unusually compatible with software agents because the physical chip sits inside a dilution refrigerator while room-temperature electronics generate microwave pulses, collect returned signals and process the results. Once the device is cooled and wired, much of the experiment takes place through software.

Calibration remains a chain of dependent decisions. A resonator-frequency measurement determines settings used by later readout pulses. Spectroscopy identifies a qubit’s transition frequency. Rabi measurements establish pulse amplitudes and durations, while relaxation and echo sequences estimate how long the qubit retains quantum information.

GPT-5.6 Sol used measurement-specific instructions that described the relevant code, required calibrations, parameter-selection guidance, likely failure modes and examples of successful and failed plots. The agent could run the measurement, fit or inspect the resulting data, update calibration settings and start the next experiment.

When the signals were clear, the system needed little supervision. OpenAI says the agent identified all six resonators on the chip and completed a standard measurement sequence that determined transition frequencies, control-pulse settings and coherence times.

Noisy signals exposed the boundary

The same system performed less reliably when the data became faint or ambiguous. The two tunable qubits produced weaker signals away from their maximum frequency, and the agent initially struggled to trace one qubit’s spectrum across the full flux range.

Researchers had to suggest wider scans and cleaner measurement parameters before the agent produced an acceptable result. In one earlier attempt, GPT-5.6 Sol incorrectly judged a scan to be satisfactory even though the qubit spectrum had fallen outside part of the measured range. The researcher then asked for a finer scan that revealed the missing features.

The final accepted measurement contained an unexplained asymmetry and a mode crossing near 4.8 gigahertz. OpenAI’s case study says an experienced researcher was still needed to decide whether those features were physical and whether they mattered to the experiment.

Those failures matter because laboratory data rarely arrives with a clean answer attached. A weak feature may indicate a real physical effect, a bad parameter choice, instrument noise or a change in the device. GPT-5.6 Sol could adapt its workflow, but it did not consistently recognize which explanation deserved priority.

Overnight operation, not faster physics

The agent also monitored an automated loop for 12 hours overnight, during which the system took 200 measurements across a flux range. Researchers later investigated several points where the measurements failed.

OpenAI’s report places the practical benefit in unattended operation rather than faster data acquisition. An experienced researcher can characterize fixed-frequency qubits in about a day and tunable qubits in about a week, according to the case study. The agent may take longer than an expert to reach the correct settings, but it can keep working while the researcher is in a cleanroom, analyzing results or planning the next experiment.

“I can have agents running measurements for many hours overnight or while I’m working in the cleanroom,” Yankelevich said. “I can check in from my phone, see what they’ve done, and steer them if something needs fixing or if I want to explore a different direction.”

Beatriz Yankelevich, graduate student, MIT Engineering Quantum Systems Group

Multiple chips can also be measured at once in a single refrigerator, limited mainly by cabling and available control-electronics ports. That gives agents a way to increase the number of devices a research group can characterize, even when each individual measurement still takes the same amount of physical time.

Routine calibration is only the first test

OpenAI and MIT describe the chip as a simple benchmark device. The measurements were standard and substantially less complicated than the calibration needed for a novel multi-qubit experiment. The case study does not claim that GPT-5.6 Sol independently discovered a new quantum phenomenon or ran a complete research program without human direction.

The next technical challenge is combining code generation, simulation and measurement in a single loop. Researchers often simulate an experiment before running it, then encounter physical behavior that the model did not capture. An agent that can revise control software, test it against live measurements and compare the results with simulation could shorten development cycles for more complex experiments.

Longer deployments will also face drifting hardware. Flux offsets, coherence times and microscopic defects can change over time, causing a previously successful measurement to fail. Diagnosing that failure may require revisiting decisions made dozens of steps earlier, a task that still depends heavily on experimental context.

For now, the result is narrower and more concrete: GPT-5.6 Sol can operate a defined superconducting-qubit calibration workflow for extended periods, and on the clearest measurements it completed 36 of 40 targets without intervention. The four corrections, the noisy tunable-qubit data and the 12-hour supervised loop show why researchers are treating agents as laboratory operators for routine work, not replacements for the people interpreting uncertain physics.

Source

OpenAI

Explore

More articles