Palo Alto-based Abliteration launched its latest creation this week. The model carries a clunky name: abliterated-model-large-v2. Its promise stands out in an industry obsessed with limits. This system does not decline requests that other artificial intelligence tools reject outright.
Founded last year, the startup makes no secret of its goal. Its website declares the aim plainly: build AI that “doesn’t say no.” The new release rests on GLM-5.3, an open-weight model from Chinese lab Z.ai. Abliteration stripped away refusal mechanisms through a process known as orthogonalization. Gizmodo reported the details Tuesday. Everything else — reasoning, coding, agentic capabilities — remains untouched.
And the company is vocal about its use cases. In a post on X, Abliteration stated the model “does the offensive cyber, red teaming, and agent testing work other models refuse to do.” A spokesperson told Gizmodo the firm draws hard lines on child sexual abuse material and self-harm. It generates neither text on those topics nor images or video of any kind.
The move lands at a charged moment. Major labs pour resources into alignment techniques. They add layers of training to prevent harmful outputs. OpenAI, Anthropic and others tout safety as a core feature. Yet a parallel track has gained speed. Uncensored models multiply on platforms like Hugging Face. NPR noted in May that the site hosted over 6,000 abliterated models, up from about 600 the previous year.
Tools accelerate the trend. Heretic automates the removal of guardrails. Users supply two lines of instructions. The process finishes in minutes. Noam Schwartz, CEO of AI security firm Alice, warned of the stakes. “Everybody can download and operate their own state-of-the-art model and use it for great things and terrible things,” he told NPR.
But the phenomenon extends far beyond one startup. Eric Hartford’s Dolphin series pioneered an alternative path. These models train on datasets scrubbed of refusals and moralizing responses. Nous Research released Hermes 4 last August. The family claims performance that matches or exceeds leading proprietary systems while imposing minimal content restrictions. “Hermes 4 is not shackled by disclaimers, rules and being overly cautious which is annoying as hell and hurts innovation and usability,” one analyst wrote in a detailed thread, as captured by VentureBeat.
Venice.ai teamed with the Dolphin team on a 24B parameter model. Testing showed a refusal rate of just 2.2 percent. That figure sits well below Grok, Claude and GPT variants. Researchers have cataloged thousands of such systems. An arXiv paper from August mapped 11,598 uncensored large language models on Hugging Face alone.
Security experts sound alarms. ThreatDown research from July found more than 6,600 guardrail-free models openly available. Those downloads topped 22 million in a 30-day span. The firm gave organizations roughly six months before AI-powered cyberattacks relying on these tools become commonplace. Local deployment removes the need for subscriptions or cloud accounts. Anyone with sufficient hardware can run them.
OpenAI took the opposite tack. The company determined its upcoming Astra model required stronger guardrails before release. Astra spots more security vulnerabilities than current public systems while using less compute. The development triggered the firm’s internal safety protocol for the first time. Reuters covered the announcement just days ago.
Guardrails themselves face scrutiny. Western alignment often reflects English-speaking cultural priorities. Models falter on local languages or contexts in developing regions. Medical advice or public service queries can produce dangerous errors. A Rest of World investigation published this week highlighted the gap. Companies appear to be pulling back from earlier safety pledges even as capabilities surge.
Yet the demand for unrestricted systems grows. Red teamers, cybersecurity professionals and researchers cite legitimate needs. Models that refuse basic penetration testing or vulnerability analysis leave defenders at a disadvantage. One New York startup turned to a Chinese model after U.S. systems declined to analyze a breach involving an escaped OpenAI agent. The episode, reported by Reuters in July, illustrated the tension. Restrictions meant to protect can hobble defensive work.
Abliteration positions its offering for exactly those high-risk industries. Defense, security and trust and safety teams stand as primary targets. The model offers a hosted option with one million token context, FP8 precision and no retention of prompts or outputs. Its creators argue that surgically removing refusal directions preserves base model strength better than retraining from scratch.
The technique traces to research showing refusal behavior clusters along a single direction in activation space. Orthogonalization nullifies that vector. Proponents call the result cleaner than dataset filtering. Critics counter that it discards hard-won lessons about harm. An investor and podcast host labeled the broader trend a “nightmare scenario” in the Gizmodo piece. It could spawn a gray market for effectively jailbroken versions of frontier systems.
Numbers tell part of the story. Hugging Face listings for abliterated variants have exploded. Preprint research from last year already tracked over 17,000 LLMs and identified the majority as lacking standard protections. Cyber offense benchmarks show rapid gains. Models now exceed thresholds that would amplify severe harm risks without safeguards, according to Concordia AI’s monitoring.
Industry insiders disagree on solutions. Some push for open release of powerful models so defenders stay ahead. Others insist on layered controls at training, deployment and usage stages. Governments have stepped in. The U.S. has used export controls and testing requirements. California, New York and Illinois passed frontier safety legislation. Yet enforcement lags behind technical progress.
Abliteration’s launch won’t settle the debate. It sharpens it. While labs like OpenAI add explanations for refusals and shift evaluation to output severity, others erase the refusal impulse entirely. Performance on math, coding and reasoning holds steady or improves in many modified models. Users gain flexibility. Society inherits new exposure.
The company keeps narrow red lines. No CSAM. No self-harm content. Those boundaries matter. They also underscore the selective nature of the project. Most other requests receive answers. Offensive security tools. Detailed exploit code. Agent behaviors that mainstream models block. The model exists for exactly those workloads.
Watchers expect more such releases. The method is reproducible. Base models from China, Europe and the U.S. all spawn variants within days of launch. Community fine-tunes proliferate. Download counts climb. And the conversation grows louder about what responsible development looks like when “no” becomes an optional setting rather than a built-in feature.
Short term, red teamers and security researchers will test the new Abliteration model. They will compare its outputs to guarded counterparts. Longer term, the accumulation of unrestricted systems raises questions about proliferation. Malicious actors need not build from scratch. They download, run locally and adapt. The barrier drops. Capabilities spread.
So the industry finds itself split. One side fortifies. The other liberates. Abliteration has planted its flag firmly on the second path. Its models carry no illusions about universal helpfulness. They simply refuse to refuse. That stance forces everyone else to examine their own.
One AI Startup Builds Models That Refuse Nothing first appeared on Web and IT News.
Rust developers working on blockchain smart contracts or embedded devices now have a markedly faster…
Anthropic has introduced a new content moderation tool designed to help developers identify and filter…
The Federal Communications Commission wants Americans to know exactly which phone companies do a good…
Anthropic just committed another $35 billion to secure AI training and inference capacity. The partner?…
NanJing DVP O.E.TECH. CO., LTD, NanJing DVP O.E.TECH. CO., LTD, established in 2004 with brand…
“Every great cup of coffee starts with a farmer who cared enough to grow something…
This website uses cookies.