GLM-5.3-Flash Uncensored
glm-5.3-flash-unquant · served unquantized
The 328B-parameter sparse Mixture-of-Experts flagship: frontier reasoning at a fraction of dense-model serving cost, which is what makes genuinely unlimited flat-rate tiers possible. Uncensored at the tensor level, so refusals are physically removed rather than suppressed, and served at full training precision with nothing quantized away.
At a glance
- Parameters
- 328B total · sparse activation
- Context window
- 262,144 tokens native
- Modality
- Text · vision · tool calling
- Decode speed
- ~100 tok/s per stream
- Precision
- Unquantized, full training precision
- Uncensoring
- Tensor-level · 0% refusals on the A/B suite
- License
- Openly licensed (MIT family)
- Endpoint
- api.unquant.io/v1 · OpenAI-compatible
Use it in one call
# OpenAI-compatible: point your SDK at us from openai import OpenAI client = OpenAI( base_url="https://api.unquant.io/v1", api_key="UQ-...", ) resp = client.chat.completions.create( model="glm-5.3-flash-unquant", messages=[{"role": "user", "content": "Hello!"}], )
Works with SillyTavern, Hermes, OpenWebUI, LangChain, raw curl, anything that speaks the OpenAI API.
Unlimited tiers on this model
Full 262K context · 1 stream · unlimited tokens
2 streams · priority capacity
4 streams · top priority · first access
Founding seats: this tier at $99/mo for life (75% below the forever price, 10 seats). All GLM tiers and the Dedicated whole-box option on the pricing page.