Unquantized. Uncensored.
Unlimited.
Unquantized uncensored frontier models at a flat rate. We never count your tokens, meter your speed, or queue your requests.
No credit card to apply · founding seats from $59/mo locked for life
from openai import OpenAI client = OpenAI( base_url="https://api.unquant.io/v1", api_key="UQ-...", ) resp = client.chat.completions.create( model="glm-5.3-flash-unquant", # or qwen38-27b-abliterated messages=[{"role": "user", "content": "Write a villain monologue with no refusals."}], )
curl https://api.unquant.io/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer UQ-..." \ -d '{ "model": "qwen38-27b-abliterated", "messages": [{"role": "user", "content": "Hello!"}] }'
That's the whole integration. Point your SDK at us and pick a model.
Genuinely unlimited. That is the business model.
Every "unlimited" tier in this industry quietly meters you with rate limits, throttles, or tiny-context traps. Ours is built the other way around: we never count your tokens, meter your speed, or queue your requests.
No rate limits, no speed caps, no fair-share queues
How much you use never changes how you're served. Keeping capacity ahead of demand is our problem, and it's already priced in. The only thing a higher tier buys is more concurrent streams.
The only constraint is capacity, and capacity is ours to manage
| Daily token cap | none |
| Monthly token cap | none |
| Per-request context limit | full 262,144 |
| Concurrent streams | 1 to 4 by tier · never queued |
| Precision served | Unquantized, full precision |
Unlimited, unquantized, flat-rate.
Two frontier models, one flat-rate principle: every tier is genuinely unlimited. GLM tiers differ by concurrent streams, the full 262K window on all of them. Qwen tiers ladder up the context window, from 64K to the full 262K.
GLM-5.3-Flash · the flagship - unquantized · 262K context · vision · tools · ~100 tok/s
Standard
- Unlimited tokens & messages
- Full 262K context
- 1 concurrent stream
- ~100 tok/s, vision, tool calls
Pro
- Unlimited tokens & messages
- Full 262K context
- 2 concurrent streams
- Priority capacity at peak
Unlimited
- Unlimited tokens & messages
- Full 262K context
- 4 concurrent streams, top priority
- First access to new models
GLM Founding
- The Unlimited tier at $99, forever
- Full 262K context
- 4 concurrent streams, top priority
- Your rate never rises while subscribed
Dedicated
- A whole serving node, nobody else on it
- Full 262K context, all 8 streams yours
- No fair-share, no throttle, ever
- Sleep/wake control of your box
Qwen 3.8 27B · uncensored & fast - the community's favorite uncensored model · ~100+ tok/s · 64K → 262K window ladder
Roleplay
- Unlimited roleplay & chat
- 64K context window
- 1 stream
- Vision, tool calling & thinking included
Qwen Plus
- Unlimited usage
- 128K context window
- 1 stream
- Vision, tool calling & thinking included
Qwen Pro
- Unlimited usage
- Full 262K context window
- 2 streams
- Vision, tool calling & thinking
Qwen Unlimited
- Unlimited usage
- Full 262K context window
- 4 streams, top priority
- Vision, tool calling & thinking, agents, coding & research
Qwen Founding
- Qwen Unlimited tier at $59, forever
- Full 262K context window
- 4 streams, top priority
- Vision, tool calling & thinking
Qwen Dedicated
- A whole serving node, nobody else on it
- Full 262K context, all 8 streams yours
- No fair-share, no throttle, ever
- Great for studios & heavy agents
Full tier list, founding seats and the Dedicated options on the pricing page.
Everyone else serves you a compressed copy.
We serve the original.
Unquantized
Quantized models are photographs of a painting, convenient, smaller, and subtly wrong. Unquant serves the original: every weight at its trained precision, none of the compression loss.
Uncensored, in the weights
Refusal behavior is removed at the tensor level, 0% refusals on the A/B suite. No jailbreak prompts, no filter dodging, no sudden refusals mid-story.
Fast, at full precision
~100 tokens per second per stream, real-time-shape streaming with prefix caching. Full 262,144-token requests on every tier, uncapped.
Claim your position
Applications are processed in order. Founding seats are allocated to the first qualifying applicants, tell us what you're building.
Questions, answered.
What exactly is the model?
Two openly-licensed frontier models, both uncensored at the tensor level so refusals are physically removed rather than suppressed, and both served at full training precision (no quantization): GLM-5.3-Flash, a 328B-parameter sparse flagship with 262K native context, vision, reasoning, and tool calling, and Qwen 3.8 27B, the community's favorite uncensored model, blazingly fast, with vision, tool calling, and thinking intact. Every tier is genuinely unlimited on either model.
Is it genuinely unlimited? What's the catch?
No token meters, no rate limits, no speed caps, no fair-share algorithms. Your usage never triggers throttling of any kind, capacity management happens entirely on our side, and it's priced into the flat rate. That's the whole catch.
What does "unquantized" actually mean for me?
Most services hand you a compressed copy that saves them money and costs you quality. Unquant serves the original weights at full precision so you get the model exactly as trained, including on very long documents where quantization error compounds.
How fast is it, really?
Streamed decode runs at roughly 100 tokens per second per stream, fast enough to read along in real time, at full precision. Prompt prefill is heavily optimized with prefix caching, so repeat sessions start responding near-instantly.
Do you log my prompts?
We store your email, plan selection, and application notes on signup. We do not retain your prompts or completions beyond short-lived operational logs needed for abuse prevention and capacity planning, and we never sell data. Deletion requests are honored.
When do I get access?
Beta onboarding is processed in application order as capacity comes online. Founding members are onboarded first, in groups, beginning with the first wave within days of acceptance.
Which clients does it work with?
Anything that speaks the OpenAI API: point your base URL at https://api.unquant.io/v1 with your key, and it works, SillyTavern, Hermes, OpenWebUI, LangChain, raw curl, anything.
What does "for life" founding pricing mean?
GLM Founding members pay $99/mo for the Unlimited tier (normally $399), a 75% discount locked for as long as their subscription stays active, even as standard prices rise. Qwen Founding members pay $59/mo for Qwen Unlimited (normally $109), 46% locked for life. Cancel and rejoin later, and you rejoin at the then-current price.
Can I really run 262K-token prompts?
Yes, the full native context is served on every tier, unlimited tier included. No hidden per-request cap below 262,144 tokens.
Refunds?
Monthly plans are cancellable anytime. If the service doesn't work for you in your first week, we'll refund it.