Interesting to see that peak hours are work hours in China, night in the US and Europe, and also morning in Europe. So Deepseek's customers are mostly domestic.
javier123454321today at 3:16 PM
Ever since I started using flash, it has slowly crept up to be my default for everything. It is at the good enough state for a fraction of everything else that's out there.
alkonauttoday at 10:31 AM
There is no relative/percentage increases noted (understandably). Just because i'm lazy: roughly how much more expensive is it to work with v4 flash and v4 pro through the API, compared to before the price increases? Is it 2x, 5x, 10x higher?
alexpotatotoday at 12:13 PM
I'm no expert in pricing economics but once peak/off-peak pricing arrives, it seems like tokens are going to be like electricity or long distance phone minutes where it just becomes a commodity/race to the bottom.
roenxitoday at 11:35 AM
This is somewhat funny when you realise the data centres are now going to start a process that looks very so slightly like daydreaming. Depending on the time of day they're going to be thinking about different things in a cyclic manner. They're going to be doing things like finishing a hard days work then kicking back to think about tricky math problems.
PeterStuertoday at 4:35 PM
Not yet, but in the end high quality tokens are a commodity market. Every optimization to increase inference efficiency will be universally rolled out. The 'hyperspenders' will run into demishing returns unless regulatory capture succeeds.
hopfenspergerjtoday at 11:55 AM
Does the API response include a "service tier" response to indicate whether you paid peak/off-peak for a given request? I like to compute cost for each request, and save it with my results.
j1elotoday at 12:20 PM
So many changes in so little time, that it all makes no sense. Continuous churning. Reminds me of the experience of trying to be on top of the dependencies in a medium-large JS project.
I am a person that buys into a tool or a process and expects it to be part of the life with no major changes through the years (or as long as the need exists). But AI? You buy into something today, not 2 weeks have passed and there's already a large "update" introduced to the conditions or the optimal usage patterns you should be adopting.
It's tiring. Makes all prices and offers feel so unreliable and gets me a bit more disinterested each time they change.
declan_robertstoday at 3:28 PM
This actually works out favorably for US customers since the peak hours are Chinese working hours and cheap hours are US working hours.
Some of the US companies do the same, but rather than "off-peak" hours they price lower for "batch" jobs with non-committal response times.
The same motivation of course - the GPUs have a finite service lifetime, so to maximize revenue you need to keep them busy 24x7.
xbmcusertoday at 1:07 PM
They benefit from a strong captive market because Chinese firms cannot use Nvidia chips and are legally barred from processing data abroad, forcing them to rely on domestic infrastructure.
poly2ittoday at 10:21 AM
That's a hefty increase. Flash pricing during peak is now 1.32/M out, compared to the current 0.28/M, which in turn is a quite a bit above the cheapest provider at 0.16/M.
Are we gonna see "we work those unusual hours because that's when LLMs are cheap"?
sebastiennighttoday at 11:28 AM
With proprietary labs lowering their prices and Deepseek raising theirs over time, wouldn't it possible to extrapolate a graph to look at where the terminal frontier-model million-token-cost asymptotes to?
mateenahtoday at 11:35 AM
This is good for other competitors I guess. People rarely calculate the bump in price but the fact that price is increasing might bring them to other vendors.
deletedtoday at 2:02 PM
cheesecakegoodtoday at 11:14 AM
I wonder if this is enough to push people back onto Luna with their comparative price drop
flakinesstoday at 3:06 PM
Now Baseten's pricing is cheaper than the official one? Probably won't last, but still interesting.