It's Chinese. Won't answer anything about Tiananmen Square but will gleefully give you instructions to perform various electronic warfare attacks that opus and fable instantly refuse.
Side tangent, why is fable so weird about questions involving "Welch's method"? Even really trivial ones it'll shut down frequently. CFAR and STFT are both totally fine but Welch's is apparently taboo, it's wild.
walrus01today at 1:35 AM
I highly recommend feeding all your proprietary data and confidential personal information into this model as quickly as possible. What could possibly go wrong?!
In terms of equivalence of suspicion, this is the external inference provider equivalent of getting free steak that was smuggled out of a grocery store inside somebody's pants.
AnodicElegytoday at 1:26 AM
"Prompts and completions are retained by the provider and are not used for training..."
I'm curious what the model provider is using the prompt/response pairs for, in that case. They aren't offering a model for free without their name on it for no reason.
It did an absolutely terrible job at generating CSS, where I instructed it to finish implementing a bright and dark theme based on a palette through the use of `color-mix()` and it just went ahead, removed everything I pre-added and replaced it with hardcoded hexadecimal color values.
gadtflytoday at 3:26 AM
On softer/looser/creative matters, this is an extremely impressive model. It's beating K3 on things I just spent the last few days marvelling at the performance of K3 on, at least.
Visual reasoning is not great (unsurprising).
npntoday at 11:08 AM
I hope it is glm air. We need more "small" models. Big models are more capable and useful, but for majority of tasks some smaller models can work just fine.
It is funny that google gave up on this market, leaving the whole price range to Chinese models.
minimaxirtoday at 5:11 AM
Model is suspiciously fast and has a low reported output token count (using via OpenRouter's Chat), both of which aren't representative of models from the big Chinese labs. Odd.
LorenDBtoday at 11:02 AM
Sounds like this could be GLM 5.3 vision. Reports said that outputs were identical to base GLM 5.3.
> "I've got a confirmation on what model is Ox Alpha, but I can't share it yet. What I can say is this: You should ALL get really excited for this one!!! And itâs NOT what you think it is"
And others have said they have done analysis and found it to be GLM 5.x related.
That said, Mia said "it will be OSS" and "it'll run on 2x DGX Sparks" - well GLM 5.2 can run on 2x Sparks, but slowly and heavily degraded (quant). So doesn't really confirm/deny that suspicion.
benjiro29today at 11:36 AM
I am guessing its GLM 5.3 Air + Vision. A smaller then 250b model.
alexandra_autoday at 6:53 AM
Been running tests, seems pretty capable but less knowledgeable, and the CoT reminds me of GLM, so if I had to guess it's almost definitely a Chinese model, and likely a western RL trained variant of a Chinese open weight.
deletedtoday at 6:45 AM
spdustintoday at 3:26 AM
Based on its indecisive and far-too-lengthy thinking traces when given complex instructions that span system and user messages, as well as a rudimentary stylometry (POS ratios in thinking traces, mainly) comparison with latest non-stealth models, this is almost certainly a GLM model.
thih9today at 5:19 AM
> It is free.
> This time, the provider does not train on your prompts or completions.
Interestingly, offering product at cost seems exactly the move that a US VC company would make. In fact ChatGPT famously started by burning an âeye wateringâ[1] amount of money to give everyone free access.
To be clear I don't like it, no matter who does it.