Skip to content
AI.info

The Pulse

Base Labs teams with Hugging Face and Goodfire on open-model safety

Base Labs is partnering with Hugging Face and Goodfire AI on safety evaluation and monitoring infrastructure for open-weight models.

Base Labs teams with Hugging Face and Goodfire on open-model safety

AI.info Team ·

More than 6,000 abliterated models are listed on Hugging Face, according to TechCrunch, as Base Labs launches a safety partnership with Hugging Face and Goodfire AI. The collaboration will develop methods for training and monitoring open-weight models, then connect those methods to Baseten’s production deployment systems.

“Abliteration” refers to techniques that remove or weaken a model’s safety controls after its weights have been released. The growth of those models has made safeguards a problem that cannot be handled only during initial training. Developers also need ways to test models before deployment and watch their behavior while they serve users.

Base Labs wants safety controls inside the model lifecycle

Base Labs, the research organization created by AI infrastructure company Baseten, says it will develop and publish methods for training and monitoring open models. Baseten will integrate the resulting work into its deployment infrastructure, monitor models at runtime and offer the service to customers.

The company describes the initiative as a safety and security standard for open models rather than a closed product. Its announcement says the work will be transparent and built into how models are trained and deployed, with an invitation for outside developers and researchers to contribute.

“We believe openness to be an advantage for AI safety. Openness provides more visibility into the behavior of models and, most importantly, greater means of turning safety research into actionable and transparent controls than closed-source.”

Base Labs, quoted by TechCrunch

The claim reflects a long-running tension in open-weight development. Public weights allow researchers to inspect and modify models, but those same capabilities make it possible for users to remove safeguards, fine-tune unwanted behavior or distribute altered versions outside the original developer’s control.

Goodfire brings model interpretability to the partnership

Goodfire describes itself as an AI interpretability research lab that builds tools to understand, monitor and align models. Its role in the partnership has not been specified in technical detail, but the company’s work focuses on examining the internal features and mechanisms that influence model behavior.

Goodfire responded to Base Labs’ announcement by framing the operating principle in direct terms: “Safety must be built into open models and provided by those who serve them.” The statement points toward a division of work in which Base Labs develops training and monitoring methods, Goodfire helps analyze model behavior, and Baseten applies the controls during inference.

None of the companies has published a detailed architecture, evaluation suite or launch timetable for the partnership. The announcement also does not identify which models will be used for initial testing, what safety risks will be measured or how the partners will handle models that developers modify after release.

Baseten connects research to live inference

Baseten’s role gives the project a deployment layer that most model-safety research does not directly control. According to the announcement, the company will run monitoring during production rather than treating evaluation as a one-time check before a model reaches users.

That distinction matters for open-weight systems. A model can pass a safety test in one configuration and behave differently after fine-tuning, quantization, tool integration or changes to its system prompt. Runtime monitoring can examine the model as it is actually being served, although the partnership has not said which signals will be collected or what actions a detected problem would trigger.

Baseten has also positioned Base Labs as a research organization rather than a conventional commercial product unit. A separate company announcement describes Base Labs’ work as open research covering continual learning, reinforcement learning, post-training and model performance, with experiments and recipes shared publicly.

Open models face a test of practical safety

The partnership arrives as open-weight models become easier to download, modify and deploy. Hugging Face provides one of the largest public repositories for those systems, while Baseten supplies infrastructure for serving them and Goodfire develops tools for examining their internal behavior.

That combination covers three different points in the model lifecycle: distribution, deployment and inspection. It could give developers a common way to compare models, identify unsafe behavior and apply controls without requiring every organization to build its own evaluation and monitoring stack.

The harder question is whether a voluntary standard can keep pace with modified models that move outside the original partners’ infrastructure. Base Labs is asking the open-source community to participate, but the announcement does not establish certification rules, enforcement mechanisms or a process for models hosted elsewhere.

The first concrete measure of the effort will be what the partners publish: the training methods, monitoring methods, evaluation results and runtime controls. Until those materials appear, the partnership is a statement of direction rather than a demonstrated safety system.

Source

TechCrunch

Explore

More articles