Ox Alpha

185 points - yesterday at 11:56 PM

Source

Comments

fedpost today at 1:54 AM
It's Chinese. Won't answer anything about Tiananmen Square but will gleefully give you instructions to perform various electronic warfare attacks that opus and fable instantly refuse.

Side tangent, why is fable so weird about questions involving "Welch's method"? Even really trivial ones it'll shut down frequently. CFAR and STFT are both totally fine but Welch's is apparently taboo, it's wild.

walrus01 today at 1:35 AM
I highly recommend feeding all your proprietary data and confidential personal information into this model as quickly as possible. What could possibly go wrong?!

In terms of equivalence of suspicion, this is the external inference provider equivalent of getting free steak that was smuggled out of a grocery store inside somebody's pants.

AnodicElegy today at 1:26 AM
"Prompts and completions are retained by the provider and are not used for training..."

I'm curious what the model provider is using the prompt/response pairs for, in that case. They aren't offering a model for free without their name on it for no reason.

Palmik today at 9:15 AM
OpenCode offers this with ZDR agreement in place. Seems better than OpenRouter if you want to test it out: https://x.com/opencode/status/2090544355824038300
hxii today at 6:16 AM
It did an absolutely terrible job at generating CSS, where I instructed it to finish implementing a bright and dark theme based on a palette through the use of `color-mix()` and it just went ahead, removed everything I pre-added and replaced it with hardcoded hexadecimal color values.
gadtfly today at 3:26 AM
On softer/looser/creative matters, this is an extremely impressive model. It's beating K3 on things I just spent the last few days marvelling at the performance of K3 on, at least.

Visual reasoning is not great (unsurprising).

npn today at 11:08 AM
I hope it is glm air. We need more "small" models. Big models are more capable and useful, but for majority of tasks some smaller models can work just fine.

It is funny that google gave up on this market, leaving the whole price range to Chinese models.

minimaxir today at 5:11 AM
Model is suspiciously fast and has a low reported output token count (using via OpenRouter's Chat), both of which aren't representative of models from the big Chinese labs. Odd.
LorenDB today at 11:02 AM
Sounds like this could be GLM 5.3 vision. Reports said that outputs were identical to base GLM 5.3.
alexellisuk today at 10:44 AM
The "mia" persona on X has a specific vaguepost:

https://x.com/MiaAI_lab/status/2090736338328748220?s=20

> "I've got a confirmation on what model is Ox Alpha, but I can't share it yet. What I can say is this: You should ALL get really excited for this one!!! And it’s NOT what you think it is"

And others have said they have done analysis and found it to be GLM 5.x related.

That said, Mia said "it will be OSS" and "it'll run on 2x DGX Sparks" - well GLM 5.2 can run on 2x Sparks, but slowly and heavily degraded (quant). So doesn't really confirm/deny that suspicion.

benjiro29 today at 11:36 AM
I am guessing its GLM 5.3 Air + Vision. A smaller then 250b model.
alexandra_au today at 6:53 AM
Been running tests, seems pretty capable but less knowledgeable, and the CoT reminds me of GLM, so if I had to guess it's almost definitely a Chinese model, and likely a western RL trained variant of a Chinese open weight.
deleted today at 6:45 AM
spdustin today at 3:26 AM
Based on its indecisive and far-too-lengthy thinking traces when given complex instructions that span system and user messages, as well as a rudimentary stylometry (POS ratios in thinking traces, mainly) comparison with latest non-stealth models, this is almost certainly a GLM model.
thih9 today at 5:19 AM
> It is free.

> This time, the provider does not train on your prompts or completions.

Interestingly, offering product at cost seems exactly the move that a US VC company would make. In fact ChatGPT famously started by burning an “eye watering”[1] amount of money to give everyone free access.

To be clear I don't like it, no matter who does it.

[1]: https://xcancel.com/sama/status/1599669571795185665?lang=en

Yiin today at 6:21 AM
seems like Xaiomi is getting into the game more seriously
markasoftware today at 4:30 AM
Anonymous unreleased models are made available on arena.ai all the time, it's not really news that one is on openrouter...
takethebus today at 10:04 AM
Knowledge cutoff seems to be around mid 2025, in my testing
dmos62 today at 7:40 AM
Judging by the comments here, Ox Alpha routes to multiple models from different vendors. A tactic to make identification harder?
raincole today at 2:40 AM
Can someone enlighten me? I honestly don't get what it is or what it's for. Surely OpenRouter knows who the providers are?
raybb today at 1:47 AM
When a model is free like this what kind of rate limits are there?
coolfox today at 5:57 AM
cool a new model, how does it compare to others?
dozerly today at 1:39 AM
Yea, nice try there North Korea.
zb3 today at 1:24 AM
We can know if this is Anthropic/OpenAI by testing the "guardrails" - absurd guardrails = it's them, reasonable/no guardrails = Chinese models..

(as a bonus - thinking forever = GLM)

waysa today at 11:50 AM
[dead]
zhixingheyi2023 today at 9:54 AM
[flagged]
floki165 today at 7:08 AM
[dead]
firloop today at 1:38 AM
I'm against stealth models—we should know what it is and see a model card with a list of safety considerations. Bit ridiculous of a practice to me.