> M5 Ultra features a massive amount of high-bandwidth unified memory, up to 512GB, and delivers a staggering 1.2TB/s of unified memory bandwidth that is 50 percent higher than M3 Ultra.
Apple never needed to participate in the AI race to zero. Because they were already at the finish line years ago building their own chips that can run large >100B parameter AI models locally.
As someone who works in AI now, I have found it pretty amazing that Apple basically didn't do much with AI software, and focused more on the hardware side. I think this is what the future of AI is going to look like, local models run on your mac for your workflow.
It's possible that they're working on their own LLM that's going to work very well on their chips, and possibly outperform anything out there when they do release it.
>I have found it pretty amazing that Apple basically didn't do much with AI software
The iPhone 15 was almost entirely marketed based upon AI (I would say fraudulently so, advertising features they still haven't delivered), and a huge portion of the OS work was on local AI or AI integration.
And for that matter Apple has been dumping enormous sums into their own AI development. Their failure to have a lot to show for it doesn't void the fact that they tried really, really hard.
It's bizarre how often this "Apple sat on the sidelines and let the AI people fight...so smart!" narrative appears on HN. Apple hasn't gone down the path of spending hundreds of billions on nvidia GPU data centres, but they absolutely tried really hard to matter in AI.
Is there anything comparable that runs Linux, doesn't necessarily look as good, but is perhaps (a lot) cheaper/fixable? Or is this really pretty optimal?
I mean this is not nvidia based right? It's all custom? So we can use it under Asahi perhaps?
I want to get something for my company to run local models, wondering what would be a good option.
I guess, what I mean is: Why are these tiny aluminum boxes so optimal?
I just want my butt ugly repairable beast machine to do the same trick. Why is my ram not unified? I have an iGPU in my server, but it can't access the 64 GB ram directly or something? It's on the CP right? Why did only Apple go for this architecture? So many questions...
well the good news is you can indeed already just have unified cpu/gpu memory on linux with an igpu on good old replaceable ram. i've done it on 8th gen intel stuff i picked up dirt cheap used. for the most part, running on the gpu wasn't faster (nor appreciably slower) than the cpu for the things i was doing, just more power efficient. overall bandwidth is relatively limited regardless which would be the bigger difference comparing against the m chips. and of course, good luck if you're hoping stuff like opencl support hasn't long ago been ripped out of whatever software you might perfectly reasonably expect to run this way today
You can't run Linux directly on these. Asahi Linux supports up to M2 only.
Linux runs very well in a VM on macOS. There are many good options for this, some free and open source (QEMU, UTM, Lima, Colima), some proprietary (VMware Fusion, Parallels).
But Linux in a VM doesn't get access to the real GPU, so model performance is limited. Those running on the CPU perform well, and those needing the GPU don't.
However, macOS on M-series macs is excellent for local models. (Maybe not as excellent as a box full of the best nVidia GPUs, but still excellent).
So if you're getting Apple hardware, like Linux, and want to run all of it locally, a fine setup for a machine to run local models, with agentic characteristics:
- macOS running one of the many local model runners. I used to use Ollama and Whisper, and now use llama.cpp instead of Ollama. Others use LM Studio, oMLX, etc. Provide HTTP endpoints to access the models.
- Linux in a VM for overall control and orchestration, with standard VM settings, and bridged networking so it appears as its own machine on your network. Also, in here provide a robust shared file server for shared state. Use this VM as your desktop and primary access to the machine, if you like Linux.
- Linux in a VM to launch ephemeral, volatile containers, with the containers using a memory-only tmpfs overlay on top of a read-only Linux filesystem in a VM disk image, with tools in this filesystem. Alternatively, a writable Linux filesystem in a VM disk image, with disk buffering set to use macOS host buffering and discard fsync requests. These settings optimise for container disk performance for data that's only ephemeral which will be deleted soon or on system shutdown. (You can combined both VMs, but need to use two VM disks to get equivalent behaviour, and be careful about VM disk configuration of the two disks.)
- Containers spawned within that second Linux VM can be spawned very quickly and run quickly, so are ideal for LLM agents that need a quick sandbox. These sandboxes generally run faster than a macOS sandbox, despite being on the same machine with VM overhead, because Linux is faster at some things. Teach the LLMs to store files and memories they want to keep in the shared file server.
It's worth noting that the 5090 (or the RTX Pro 6000 big brother with 92GB VRAM) will run rings around the Mac when it comes to compute.
My old 3090 is typically significantly faster (almost 2x token/s) than my M4 Max 128GB machine, as long as the model fits in the 24GB of VRAM.
In most situations it's a better idea to just buy tokens. But there are definitely cases when that's not an option. And then a machine like the M5 Ultra can allow you to do things locally for a fairly limited budget. And in a simpler package to manage than a machine with multiple GPUs.
I was so blown away at all the discourse surrounding "Apple fumbling on models". They should never have been in the model game to begin with. Apple crushes hardware over the last decade and that's a huge advantage today. In the end, massive models have proven to be very strong, but small models have proven to be good enough (especially with the recent Qwen 2.8 27B drop) and that's where I imagine the future will lie for consumers.
They did participate early on with Apple Intelligence and failed miserably. Really good move to not double down and let the others explore the space first
There are some older games (not that old, one of the metal gear/raiden games had it, so like 10-15 years) where the cut scenes are actually prerendered video playback - so not very demanding. Don't know if mixtape does this though.
Apple Studio with maxed out M5 Ultra, 256GB RAM and 16TB storage is 18,299$. The 512GB RAM version apparently is coming in October, considering that the difference between 96GB and 256GB is priced at 4000$, the 512GB upgrade must be eye watering.
So, on the mini the RAM upgrade runs at 25$ per GB on all tiers, the same as the Studio therefore the upgrade to 512 will probably cost 6400$.
The fully maxed out Apple Studio then will be 24699$. It's 17199 if you don't upgrade the storage.
So is downpayment on a house. I would buy the house and just pay for tokens as needed. The house will get more valuable and that wealth would buy a lot of tokens in the future - which will probably get cheaper.
EDIT: or buy AAPL. If I had bought Apple stock instead of buying a Mac LC II in 1992, then I would have about $2 million in Apple stock.
Or more simply – $25K (+ tax) put in a savings account will earn about enough interest to pay for a $100/month AI subscription indefinitely. And at the end of it you still have the $25K.
Currently m5 max has a prefill rate for Qwen 3.8 27B of 400+ tok/s (and it can be improved further, the software is not yet at the level it is for CUDA)
That means for context under ~5k or so ttft (time to first token) it's going to respond faster than Claude. If the answer is less than ~1k I think the request finishes sooner. And it's ~claude 4.5 or 4.6 level intelligence.
It’s basically impossible to compete on economic terms with deeply subsidized hardware that is widely available to rent or as a service with zero commitment.
For general inference there’s no ROI that makes this work vs subscriptions.
25k for computer now, plus 9-10% sales tax, plus operating cost, plus time and cost for R&D tinkering with models, harnesses, and infra (assuming highly capable engineering talent that can get paid for your human inference) vs a HEAVILY subsidized subscription at 200 per month with free R&D has a pretty long ROI (15 years?)
At API costs, it’s like 6 months if you’re heavy on inference.
For training, specialized models will have their own ROI that makes this worthwhile. Then debate renting capacity and the platform to choose
Yeah I think you call out the right incentives. Lots of people are very interested in renting out the hardware, and it gives us lots of flexibility vs buying.
If the pace of hardware change is what the jensens of the world say, then the rents on existing hardware will decline.
Why are we comparing to the maxed out model only? The M5 Ultra with 96 GB, 1 TB, and 64 GPU cores is $5499.
Apple lets you lease that model for $110/month for 36 months.
If you can settle for a M5 Max base model that would be a $49/month lease.
Today, you should be able to run Qwen 3.8 - 27B amazingly well on either which is giving comparable performance to 5.6 Luna on SWE Bench. The local models are now getting better and more efficient and this should give you headroom. Tools like turbo fieldfare are really reducing the memory requirements to run large models and I don’t see it stopping soon.
Maybe they had the cash on hand, maybe they didn’t. Assuming they are in the family formation stage of life with young kids, an outlier risk appetite would be required to sleep well at night.
Well, this is all assuming a 20% down payment, which anecdotally as a 26 yr. old, nobody I know is able to achieve. All my home-owning friends put 3-7% down. Granted, most are using first-time homebuyer loans which are generally more favorable.
Depends on where you live. Also keep in mind population decline so long term, im not so sure. Prices already dropping in less desirable places (everywhere outside of most blue areas) (Assuming you are talking about the US sorry if not)
Apple hasn't been selling just ram for a long time, they sell vram. Try getting 512 gb of HBM on current Nvidia cards - it's gonna cost way more than $ 24k. And here you get the same amount of memory for weights right in a quiet unit under your desk
Put together a similar build with a couple of rtx 6000 Ada cards and Apple's price tag suddenly looks pretty damn reasonable
HBM itself is very expensive but it’s not really fair to compare to LPDDR or GDDR
They’re very different things.
The more logical argument to me is that Apple uses its upgrade price points as more than just direct BOM and rather as a proxy for things that are amortized across all their sales like support/warranty/etc so higher SKUs subsidize the costs of the lower ones.
i remember when SGI boxes were $50k and then literally worthless just a couple bears later. i remember my university had a pile of them for free outside the deans office.
"Nobody" puts down 20%. That hasn't been a requirement for decades (in the US).
And there are plenty of places in the US where you can buy a house for $150-$200k. Maybe not places you want to live for one reason or another, but they're there.
Ten years ago a bought an expensive MBP because I do a lot of stats in R, Python etc that benefited from it. But the next Mac I’ll buy will be a much lower-end model, because it’s just easier these days to do that work in notebooks in the cloud.
I don't think 16TB storage is a right choice. Going 2TB and it's 11,299. Probably, if you buy the storage and install it yourself you can go higher and quite cheaper.
I think the reason they offer these options in the first place is the discontinuation of the MacPro and their remaining need to offer high end solutions.
Regarding the $25 price per GB RAM. These chips use LPDDR5X RAM. Looking at the Framework site, they are selling LPDDR5X LPCAMM2 modules for the following prices:
I recently had to buy RAM for some on-prem servers at work. 256GB was around $3800, and it wasn't even particularly fast RAM... it also wasn't ECC, because that would have spiked the price still higher.
Maybe in a year or so or maybe never. It depends on how 14a turns out and if it is comparable to TSMC 2NM. They may also choose to utilize 18a-p / 14a for other chips and not the M-series.
The 512GB Ultra is amazing, sure. But who is it for exactly? VC funded big spender founders? In that case why would they need local AI? The ultra rich enthusiast? But there can't be too many of those. So who actually buys these?
this comment has always existed behind every apple release, most especially anything vaguely pro-ish.
to answer your question : looking at the aftermarket availability of Apple's prior best and brightest : practically no one buys them.
"people here buy them" , well, 'here' is one of the most affluent groups of people in the world.
They're available as movie and television set pieces (undoubtedly disappearing into the home of someone close to the staff post-production), and for administrative/boss types that can slip the cost into a ledger somewhere that few will ever see.
It has been a hobby of mine every few years to check out the apple site and see how big I can option a machine. My record was when I was in high school years ago and was able to option some pro studio-ish apple desktop thing to like 61,000 usd out the door.
Local AI is the future, and a lot of people want the first mover advantage or to toy around with it. I know a guy with a small rack of Nvidia Spark machines that he uses for that purpose; it's as much as a decent used car.
The competition is a custom multi-GPU NVIDIA RTX pro desktop, which go for much much more. $20k is cheap for 512GB addressable memory. The old mac pro could easily be configured to cost that much.
If I can get Sol level capabilities on a $20k machine, then it is well worth it for my employer to buy me that machine for work as a workstation. When you start paying in tokens vs subscription costs due to enterprise agreements, you really start to see how much cash utilizing frontier models at the frontier costs (and I'm efficiently using luna and other models where possible!)
Yeah, I'm not sure people realize how expensive ZDR/Zero Data Retention is, and how important it is to a lot of businesses, this kind of thing starts looking really cheap really fast if it's a reasonable substitute.
Even using multiple windows in parallel for as many as 5-10 hours per day, I find that I am not fully using my claude max (20x) and chatgpt pro (20x) accounts. I can for sure use up the claude max account, but chatgpt either gives me a free reset before I run out of tokens or I just fail to use the full quota. The quota for Sol seems like 10x that of Claude Opus at the same level, and forget Fable, you can use a 5 hour quota in 20 minutes.
But lets do the math:
Lets say a 20k workstation can run 1 inference at a time at the same speed you get with Sol hosted by openai (big assumption) and run an equally capable model (big assumption).
Each month this gives you about 100-170 inference hours on a Sol 20x Pro account, and 720 hours (if you utilize 24/7) on the workstation.
Assuming a 36 month amortization before the workstation has to be replaced due to no longer being able to run frontier models or is too inefficient due to electrical costs or what have you:
The monthly workstation cost is about $550 capex and $150 electricity -> $700/month
You would need about 6 Pro accounts to reach that capacity, which would cost you $1200 a month.
But this fails because:
- You most likely can't utilize the workstation 24/7. Your work hours will be concentrated into 6-10 hours per day.
- During work hours you are capable of utilizing more than 1 concurrent session. 6 Sol accounts would support as many as 20-30 during working hours, not all the time but if you could burst to that many (don't forget sub-agents and agent directed parallel agent workloads).
- In 1 year the cost of Sol level models is likely to cost a fraction of what it does now.
this leads to:
Workstation 1 Sol Pro 2 Sol Pro
Monthly cost $700 $200 $400
Raw capacity (hrs) 720 120 240
Usable capacity (hrs) 100-130 120 240
Concurrent sessions 1 3-5 6-10
$ per usable hour ~$6.00 $1.67 $1.67
Usable hours per $700 ~115 ~420 ~420
Subscriptions are, and will likely remain, the best deal in town. Unfortunately, larger companies aren't able to do that. When your monthly token costs are in the $5-10k range, the local inference starts to look a lot more attractive
In the case where you pay for tokens without a subscription, the analysis is still very much not in favor of buying hardware.
The assumption previously used was that you can run a Sol level model on an M6 or whatever hardware $20k gives you. That is not true, it was an assumption made to show that even giving your own hardware every reasonable advantage it still loses.
Lets compare buying tokens of the best model you might run on your own hardware (still being unrealistic in favor of your own hardware) vs that same class of model on the market. I think one of the best models you might be able to run is GLM 5.4, but lets just look at chinese models generally:
$20k workstation, best case: $15k M5 Ultra 512GB, 36-month amortization, ~$440/mo. Runs a GLM-5.3-class model at ~30 tok/s. Saturated 24/7 it produces roughly 58M output tokens/month.
Buying those tokens:
DeepSeek V4 Pro @ $0.87/M $50
Kimi K2.6 @ $4.00/M $232
GLM-5.3 @ $4.40/M $255
Kimi K3 @ $15.00/M $870 (does not fit on the box)
So yes, at that speed for sure. But if the speed goes up? or the ability to batch at the same speed goes up? The economics start to shift. The gap is much closer, and you'd end up with a box you can still use or sell later.
As speed goes up the cost / Mtoken will necessarily go down at roughly the same ratio so it will wash out. The still use hardware or sell hardware value is factored in to the amortized monthly cost, it assumes a 3 year markdown, and does not factor in the cost of money which should almost cancel the resale value in the end, which I think is quite accurate (any residual cost on a graphics card after 3 years is so small compared to the current price it should be discounted and in included there).
Where you might win by owning your own hardware:
- Hardware costs go up, and thus api costs go up. You've locked in your pricing.
- Chinese/Open models become illegal/hard to access the way we do now. OpenAI and Anthropic are trying very hard to build a regulatory capture scheme to do this. I think they will be unsuccessful because China just won't participate.
> $20k workstation, best case: $15k M5 Ultra 512GB, 36-month amortization, ~$440/mo. Runs a GLM-5.3-class model at ~30 tok/s. Saturated 24/7 it produces roughly 58M output tokens/month.
For agentic coding, ~90% of the cost comes from cached input tokens. This cost increases quadratically with the session length. If sessions go near 1M context, the number of cached input tokens can easily exceed 1B in a day.
> - You most likely can't utilize the workstation 24/7. Your work hours will be concentrated into 6-10 hours per day.
isn't the whole point of all this ..... agents? isn't that what literally everyone is always clammering about in these threads? in which case the workstation is useful 720 hours out of 720 hours.
"You most likely can't utilize the workstation 24/7. Your work hours will be concentrated into 6-10 hours per day."
I have agents running 24/7 doing research, in fact I would argue this how they will be used for most programming tasks in the near future. For chatting, I agree local inference makes no sense. But for tasks that run continually, I'm not so sure. Personal computers took a while, local inference will too, but I think it will happen.
Sorry for the snark, but are you trying to cure cancer? What could possibly need 24/7 research in our domain, that doesn’t need your input every 30 minutes?
Normal boring CS scientific work. Just running running my experiments, reproducing other papers, etc. A lot of it does involve the agent waiting for some computation, but the fact that it resumes independently when I'm sleeping is kind of the point (+ usually I have several running in parallel).
I'm not trying to cure cancer, although I do hope people who are use LLMs. ;)
One of the advantages of LLM's is that you can set up a task list and tell it to burn through those tasks overnight. It's a rare night that Claude isn't busy for me, and I do burn through my 20x subscription, sufficiently that I downgrade from fable around about ... now in the week...
Other than the local AI crowd which is much recent it is professionals using Final Cut Pro for video editing, Logic Pro as a DAW and music production, Video transcoding, Photoshop and other tasks for high performance computing that don't need Laptops but want above 128GB of ram and prefer a Mac. Then there is the obvious group of developers that are making Apps for all of their products. Also, these are great for the workplace. AI is much more recent thing that Apple products were used for.
If Apple didn't sold these things they wouldn't make them but, also the level of marketing that Apple is talking about for AI is basically the new group they need to capture because the ones I just listed are already buying Macs and or easily to motivate with the other obvious CPU / GPU performance upgrades for code compilation, faster memory and video transcoding.
I think "ultra rich enthusiast" is in the right ballpark. There are people betting on being able to create their own revenue generating products and services with their own local hardware and very little operating costs. That may or may not make sense as a business idea. But people with wealth and risk appetite trying a new kind of business model and cost structure has a strong tradition.
Put another way: If $25k is the full extent of the start up capital costs, and operating costs are very low, that is a much cheaper business to start than most! The question is whether this is actually a useful model for a revenue generating business. I think that remains to be seen.
My guess is that there will be a few hits (which we'll hear a lot about - especially when someone actually pulls off "the first single-person unicorn", which I do suspect will happen someday) and a huuuge number of misses, which we won't hear much about.
Enthusiasts buying these for fun are not the target market. These aren’t big sellers to begin with but a lot of the sales are going to companies where people have budgets for gear like this and can make a business case for it.
This is, sadly, probably a foreign concept to a lot of people who have only worked at companies where hardware purchases are viewed as something to minimize and everyone is stuck with the same low spec laptops that the finance department picked out. At companies where someone might have a legitimate use for a $20K machine, their fully loaded costs (not their salary) are $300K or more, and other teams like sales are spending thousands of dollars per week on things like travel and hotels for their job, spending $20K on a computer that’s going to last several years is not a hard choice.
My friend plans to get the 512GiB M5 Ultra ASAP. He makes money by selling synthetic data and training models for companies in San Francisco and it's the cheapest way to be able to do that on your own local hardware. It's also a bit irrational as he could save money by just renting, but I get the appeal of owning something outright.
I realize it's not a 512, but for context on why I ordered a 96GB. $95/month after a trade in. I'll use it for a long term project where I'm building a national sized dataset/processing video transcripts with a rubric (of churches/sermon health). I'm willing to trade SOTA models for local ones for cost of regenerating/model consistency/ethical reasons.
I'll find other ways to use the power though, opencode or maybe start working on more video projects.
Anyone dealing with sensitive or regulated data can get started instantly with local models where they may need time consuming process or complexities to send sensitive data outside.
Aside from high-spending users, I think many people use it as a productivity tool. If you can use it to make money, and the money you earn far exceeds the monthly payment, why not make your work smoother?
I use both the base M4 Mac mini and an M4 MacBook Air for Final Cut Pro, Photoshop, and Fusion, and while they're not as almost-always-perfectly-smooth as my Mac Studio, they're about 100x better than the experience I used to have on my old Intel MacBook Pros in the 2010s.
I edit 4K ProRes and H.265 footage, sometimes with multicam (up to 4 streams) and color adjustments, titles, etc. It's only after stacking 3-5 effects before things can stutter, really.
Or if you try doing something CPU-intense in the background _while_ running some heavy creative software. I just don't do that.
Even w/the pricing spike, inflation adjusted we are back to roughly the prices of a new Mac SE/30 for something that can beat a turing test w/o sweating.
I yield the floor to no one when it comes to pessimism, but that's incredible.
Yeah...but in context, "gullible" seem a bit pejorative. Humans are also hopelessly incapable of sensing radioactivity, methanol in their alcoholic drinks, carbon monoxide, and a great many other things that our ancestors just didn't encounter much.
Though we're pretty good at sizing up a person's emotional balance/maturity and competence at familiar tasks. So maybe have an old blacksmith watch the AI/robot interact with horse owners for a while, then shoe their horses, and see how well it does.
One needs to be a little more sophisticated about it. Ray Kurzweil, for example, set fairly clear rules for his interpretation of the test in his 2001 wager:
It depends who takes the test. I am not yet, to my knowledge, fooled by AI.
I've tried [1] and I almost 100% detect which is the AI. I really want to convince myself I have failed, does anyone know of a better site/resource for this?
I know it might be moving goalposts but I would consider AI to have passed in a well and truly undisputed manner when [2] is resolved.
But in a more practical sense, if AI can impersonate humans so well today then why are state of the art frontier models so obviously AI when they create PRs, commit messages, documentation, etc. Are the companies deliberately making them unnatural?
we might need to bring back the Voight-Kampff test. anthropic at the very least is introducing a water making system to Claude which might make them more identifiable to humans as well as much easier to detect for machines.
The Turing test is more complex than what gets suggested.
An the "popularized" version is faulty also since it uses an ideal, abstract human judge (like the "spheroidal economic agent").
But if you want to add declinations to the said popularized image of the Turing test, you may add Maxim Lott's IQ tests at trackingai.org . Around January 2025 LLMs reached an equivalent IQ of 100, for example.
I built a new PC about two years ago, and I probably got it at the last possible opportunity for a while. CPU and motherboard have come down by maybe £100 in between the two of them, but a 7900 xtx (or any other 24GB GPU) for under £1,000 now seems like a bargain, and £180 for 64GB of DDR5 makes me feel like an old man talking about the halcyon days.
I bought a motherboard from Newegg in late 2024. It came with 16GB of RAM as a free "gift". The same RAM goes for $534 on Newegg today. It was $437 in January. Prices are still creeping up for this "throwaway" memory.
I still have it, but never installed it because I only use ECC memory in my desktops. I'm saving it for a future Desktop in case memory prices never fall again.
Welcome to the club. If you're _really_ competitive in cs2, I'd swap out to a 9800x3d setup, but it's still a maybe. Very little reason to upgrade right now other than to run LLMs.
I for one can't wait. Current prices are absolutely insane.
However a problem exists with scalper bots. They're going to fight tooth and nail to keep prices artificially high even when supply increases. Same with the hyperscalers blowing "free money" at eating supply as a strategic position against smaller competitors.
That sufficient supply is maybe on the distant horizon. In the meantime everyone is being fleeced because the scalpers have the only supply available. Even if there's a supply glut tomorrow the scalpers have the infrastructure in place to buy out retail channels keeping prices high.
That's an extreme risk to them buying at inflated retail prices when supply is improving and prices are about to come down. Scalpers arent a large enough group to act as a cartel.
Cycles tend to be 5-10 years long not 25. Without knowing anything else I would expect a fab you seriously start planning today will be at full capacity in about 5 years. Nobody serious likes delays - in particular the banks don't like loaning money that won't at least start paying off. They know it takes some time to design a building - but factories typically are standard buildings so once you know about the size you can get it done fast - I expect 1 year to have the building done is the worst case (and it can be done in 3 months possibly if your project management is good - after interest this is cheaper than the 1 year). It takes time to build and install the specialized machines that go inside - this is the largest problem, but you typically order them first and then plan the building around the needed space and when they will arrive. Then you need 6 months to setup the inside of the building. From there it is just ramp up time.
The above is a standard project management problem. We do this for lots of industry all the time. There is every reason to think you can get a new factory running in 5 years.
Note that I said 1 factory above. Some of the special machines we don't have the ability to make them fast enough to do 2 (I don't know the real number!) new factories in 5 years. Existing factories are using most of the special machine capacity to replace machines that wore out on the way - this can be corrected as well, but it adds another year and the expenses are much larger. Realistically though 1 new factory is likely enough.
I'm expecting it to somewhat collapse. I don't know if it'll go back to pre bubble prices (here's to hoping), but I do expect a pretty sharp decline around 2030... probably not before then.
Basically everyone that makes memory is building new fabs, meanwhile I'm not sure how much longer AI datacenter demand for ram will last. I think the decrease in AI ram demand and the new fabs will likely coincide leading to a collapse in pricing.
That is, of course, assuming the memory manufacturers don't pull their favorite trick and collude.
If you don't understand that the AI driven memory boom has totally broken the cycle you are going to lose a lot of money. There is infinite demand for Intelligence and that translates directly to chips.
Absolutely, I am in a hold pattern for the next few years. I suspect in another 3-4 years you will be able to build an absolutely stacked machine at a reasonable price. Probably won't go back to what was before but it will be decent.
Sure, but this thread started with somebody pointing out that these computer prices are in line with Mac SE/30 prices. I'm just carrying that over to cover storage and it's even more amazing how much we get for so little money.
I happily booted and ran an Intel Mac mini using a 4TB drive in a Thunderbolt 3 enclosure, and I do the same for an M4 Max Mac Studio using an 8TB drive in a USB4v2 enclosure (OWC Express 1M2 80G).
You'll just have to be careful to match the enclosure to the ports on the system. The base-model M6 Mac mini still uses Thunderbolt 4, so a USB4v2 enclosure would be wasted.
Funnily enough I do the same with my Intel iMac... I have a 2TB Thunderbolt 3 Glyph drive as my primary and boot drive and it works a dream... if I got one of these new Mac minis without upgrading the drive, presumably I could just plug this straight into it and apple will do its magical migration thing?
These Macs are reportedly shipping with macOS 27 Golden Gate. You'd have two incompatible constraints: Macs don't like to run a version older than what they shipped with, and Golden Gate doesn't support the Intel iMac.
You might be able to boot from the internal drive, install/upgrade the external drive, then boot from it.
There is nothing special about making a computer that's unaffordable, irrespective of whether it's 2026 or 1989. If anything, it just shows how little they're trying.
I live right next to a micro center and remember when they started offering that deal... still so pissed at myself for not buying one. I ended up just buying a raspberry Pi for what I was doing, but seeing as where the prices are now, I messed that up a bit. Also my worst sin was not buying 64GB of DDR5 when I was doing my computer upgrades back in August last year.
I returned an M4 Mac mini, 64GB, unopened... because I thought it was excessive for my needs then. I swear it'll be one of the things flashing before my eyes when this all ends.
Out of all of my home servers, the M4 Mac Mini is probably my favorite and runs many workloads even though it sits on a 96 core Epyc server now (needed the massive amounts of PCIe for something else). I was able to upgrade my SSD with a 3rd party one, not that sticking an NVMe drive on one of the thunderbolt ports would have been a bad pick. About the only place it didn't make sense is if you wanted a ton of RAM.
I feel so incredibly lucky I chose to overdo it a bit on RAM when I upgraded last year... though I remember deliberating to go for 192GB (it's 96GB now) but motherboard support was somewhat more complicated and it felt like a waste of 300 euros...
Yeah. One place I worked went down under and they offered to hand out on-prem servers for free. These were 192 GB DDR4 each. I was moving houses and lazy, so I declined. I still think what would've happened if I just rented a car and went downtown.
Sick. Particularly stoked for the 10gb network card ($100 option) when using the Mac Mini as a server. Just wish the memory + NVMe prices could come back down to pre ai-goldrush prices. As $2999 for the M5 Pro with 64GB RAM feels painfully over-priced.
My example is I'm looking for a new media creation machine to replace my homebuilt PC from 2015. Since that old machine can't run Windows 11 and also because Apple storage is so expensive, my idea is to turn the PC into a Linux storage machine with a 10G NIC. Then I should just be able to edit off of the storage instead of worrying about caching it locally.
The prices are ridiculous though. I may just keep rolling with my Windows 10 setup.
As long as everything is using 10gb, faster transfer speeds when moving data between computers. For downloads, it wont help unless you have 10gb fiber, but for most folks, 2.5gb is quite fast. Hell my spinning media NAS has a hard time saturating when moving files between internal servers.
RAM and SSD in apple gear has always been way over-priced. There was a short blessed period in March where the M5 Max macbook pro was out, but the general 30% price hike had not yet happened. In this period, given the insane inflated RAM prices, the price apple was charging for the M5 Max with 128 GB RAM was actually _reasonable_.
M5 Pro in a Mac Mini with 64GB RAM and 10Gbit Ethernet seems like the perfect Jellyfin server and Ollama test server. All for just over $3K (I specced with only 1TB local nvme).
The page says 170G/s memory bandwidth for the NPU and 1.2T/s for the GPU. Why the discrepancy if it's all "unified memory"? The former is nothing to write home about as far as AI compute is. The latter is really nice.
Which one is it you can run local models on? I suppose the NPU only.
Apple still has the best hardware so I moved to it for the last few years, but the closed software ecosystem is terrible for taking advantage of it.
I wasn't able to debug network errors (restartin my Mac worked), Metal was missing low level disassembly / debugging tools (there is some hard to use UI), but the worst thing was the inflexible windowing system.
Even getting all the window handles on all screens/desktops with their titles and programs is impossible.
I just decided that I move to Omarchy 4 (basically Hyperland + QuickShell) + NVIDIA GPU, and I already was able to customize it more than my Mac in years.
I will miss Apple's hardware for sure, but not MacOS and the missing hardware documentation
I had a chance to try Omarchy past few days and it's just very different vs macOS. I honestly never had a problem with the windowing system ever since I built my own customization scripts (i.e. Hammerspoon). I can see the appeal for someone who wants ultimate customization though but macOS still wins overwhelmingly when it comes to polish, ecosystem, user experience, and ecosystem (nothing comes close to macOS apps).
It’s not just Omarchy, there’s really not much out there in the desktop Linux sphere for those who are mostly happy with how macOS works out of the box. Everything is either in a similar vein to the Omarchy setup (hyper-minimal tiling WM), Windows-like (KDE, Cinnamon, most other DEs), or a chimera with a grab bag of design bits from every desktop and mobile platform (GNOME, Pantheon, COSMIC).
It’s a bit depressing because it means that if I ever feel forced to switch my daily driver, it won’t come without a dump truck load of friction, frustration, and lost productivity, which I’ve validated by using the various Linux desktops on secondary machines.
There were 1000 plugins created for Omarchy 4 in 2 days. That's why I don't feel it being hyper minimal anymore.
It's still not well integrated of course as those plugins are from different people, but I at least don't feel powerless as I know I can make any change easily.
Perhaps, but one of desktop Linux's defining philosophies is the machine working for and adapting to the user rather than the reverse, so it's a letdown that this only really applies if you're coming from Windows or have been a Linux user all along.
Choices are always opinionated by the dev but you can always fork and change to obtain what you need. MacOS doesn't let you do that and isn't easier to switch to.
With the caveat that one has the time and energy to fork and modify in the first place, and then additionally has time and energy to maintain the changes that upstream won’t take…
You can't expect any system, even the most configurable ones, to be ready out of the box to everyone's particular nitpicking without a minimum of effort.
I run KDE on a secondary machine and it's the best it's ever been, but for me it's best for single-purpose machines where desktop differences basically don't matter. It's about on par with one of the better Windows releases (XP or 7 maybe) in terms of how frustrating I'd find it to do work under.
It's interesting because I just haven't felt the polish.
For example when using PyTorch I wanted to try to speed up my NN kernel by 2x by just using half precision and haven't noticed any speedup at all.
My main program missing from going back to Linux was ChatGPT Desktop, but now it's there.
I just checked out Hammerspoon, I'm happy for you that you wrote it, and looks great, but it has the same problem that I had: for security reasons Apple stopped allowing the window APIs to get all important information on other workspaces. You can only do it with Accessibility API. I was trying to fight with it but have up.
If the browser being Chromium-based isn’t a hard requirement, it may be worth checking out the Firefox-based Zen Browser[0]. Its UI is very similar to that of Arc, to the point that I’d call it Arc’s spiritual successor.
How are you liking Omarchy? I saw a video on it recently, and it looks 'pretty' but still looks like it's a lot of memorization of shortcuts and feels like the 40% keyboard of OS's. Like some people it's absolutely amazing, but lets be honest, it's going to be really difficult to be as productive as a full fat keyboard.
As a long time Emacs user, I beg to differ. Muscle memory will settle in after a few days. I still use Emacs shortcuts to navigate text boxes on macOS.
The most disruptive thing in Omarchy is not the keyboard, it's the tiling WM in my opinion.
There's aren't really any alternative OS [1], but I don't think it's fair to say there's a closed software ecosystem. You can install/compile/run whatever you want, including kernel extensions. The hardware isn't locked.
> Apple allows booting unsigned/custom kernels on Apple Silicon Macs without a jailbreak! This isn’t a hack or an omission, but an actual feature that Apple built into these devices.
I wanted to try it but it works just on M1 and I have M4. I was also considering upgrading to M5/M6, again to take advantage of what I really love: access to the 2nm process.
The performance you get with M5/M6 in theory is crazy (especially having matmul cores with such an energy envelope and unified memory). By the time I can use it well in Linux there will be M7/M8 though.
With NVIDIA I can buy a Blackwell based core right now and use it (even if I will have much less GPU memory. Though it will be high bandwidth).
It’s barely a “feature” that they let you execute arbitrary code on your own hardware.
They don’t provide any source code or documentation (which they easily could) so if want to do literally anything with your “open” device you must first reverse engineer the entire MacOS driver stack like the Asahi project did. Great fun, but it’s a criminal waste of talented human effort.
On top of that, the reason Asahi doesn’t work on M4 and later is that Apple has intentionally modified the ARM core to prevent stock MacOS from running once the hardware is “unlocked,” which makes it way more difficult to reverse engineer.
Give IBM some time, and there will likely be similar stuff on Linux distributions. The Linux kernel has been dropping hardware support, and systemd has been getting some... interesting features.
Naturally, the benefit for Linux is still the variety available in distributions. While IBM's RHEL and Canonical's Ubuntu may adopt some questionable stuff, Slackware, Devuan, Omarchy, and such likely won't.
Features to replace installation, cloud metadata uniformity, better secure boot support, boot loader support, and improved TPM support. None of this sounds user hostile, but taken together and handed to the corporate Linux vendors? This easily becomes a way to make Linux a licensed appliance. Yet, handed to a group like Arch, it just enables wildly cool stuff. Incentives are a thing.
I don't understand how absolutely basic things are not possible on Mac OS. Switching between two windows easily, displaying hidden files in the finder...
These things would be way easier to solve than building the next generation hardware.
I think the worst is the lack of support for 'click-through' behavior on MacOS. This is that you have to select a window first with left-click before being able to select UI elements within that window with left-click. The amount of time lost with extra clicks is incredible.
It is even more frustrating that this behavior is somewhat inconsistent, with some applications allowing this and others not (although it is more often the latter).
The clickthrough behaviour used to be the default in NeXTstep (and it was one of the things which I misliked not having when switching from my Cube to my work Mac and ThinkPad portable ages ago).
The apps it works in are probably written in Cocoa née "Yellow Box" (allegedly so named because Bill Gates stated that rather than write apps for NeXT APIs he would instead....)
The way I fix it when using my Mac these days is to only use apps written using Objective-C and so forth.
1,025 comments
[ 0.23 ms ] story [ 155 ms ] threadApple never needed to participate in the AI race to zero. Because they were already at the finish line years ago building their own chips that can run large >100B parameter AI models locally.
It's possible that they're working on their own LLM that's going to work very well on their chips, and possibly outperform anything out there when they do release it.
The iPhone 15 was almost entirely marketed based upon AI (I would say fraudulently so, advertising features they still haven't delivered), and a huge portion of the OS work was on local AI or AI integration.
And for that matter Apple has been dumping enormous sums into their own AI development. Their failure to have a lot to show for it doesn't void the fact that they tried really, really hard.
It's bizarre how often this "Apple sat on the sidelines and let the AI people fight...so smart!" narrative appears on HN. Apple hasn't gone down the path of spending hundreds of billions on nvidia GPU data centres, but they absolutely tried really hard to matter in AI.
I mean this is not nvidia based right? It's all custom? So we can use it under Asahi perhaps?
I want to get something for my company to run local models, wondering what would be a good option.
I just want my butt ugly repairable beast machine to do the same trick. Why is my ram not unified? I have an iGPU in my server, but it can't access the 64 GB ram directly or something? It's on the CP right? Why did only Apple go for this architecture? So many questions...
Linux runs very well in a VM on macOS. There are many good options for this, some free and open source (QEMU, UTM, Lima, Colima), some proprietary (VMware Fusion, Parallels).
But Linux in a VM doesn't get access to the real GPU, so model performance is limited. Those running on the CPU perform well, and those needing the GPU don't.
However, macOS on M-series macs is excellent for local models. (Maybe not as excellent as a box full of the best nVidia GPUs, but still excellent).
So if you're getting Apple hardware, like Linux, and want to run all of it locally, a fine setup for a machine to run local models, with agentic characteristics:
- macOS running one of the many local model runners. I used to use Ollama and Whisper, and now use llama.cpp instead of Ollama. Others use LM Studio, oMLX, etc. Provide HTTP endpoints to access the models.
- Linux in a VM for overall control and orchestration, with standard VM settings, and bridged networking so it appears as its own machine on your network. Also, in here provide a robust shared file server for shared state. Use this VM as your desktop and primary access to the machine, if you like Linux.
- Linux in a VM to launch ephemeral, volatile containers, with the containers using a memory-only tmpfs overlay on top of a read-only Linux filesystem in a VM disk image, with tools in this filesystem. Alternatively, a writable Linux filesystem in a VM disk image, with disk buffering set to use macOS host buffering and discard fsync requests. These settings optimise for container disk performance for data that's only ephemeral which will be deleted soon or on system shutdown. (You can combined both VMs, but need to use two VM disks to get equivalent behaviour, and be careful about VM disk configuration of the two disks.)
- Containers spawned within that second Linux VM can be spawned very quickly and run quickly, so are ideal for LLM agents that need a quick sandbox. These sandboxes generally run faster than a macOS sandbox, despite being on the same machine with VM overhead, because Linux is faster at some things. Teach the LLMs to store files and memories they want to keep in the shared file server.
But you get a generic computer and much more RAM.
And you lose a couple of organs.
My old 3090 is typically significantly faster (almost 2x token/s) than my M4 Max 128GB machine, as long as the model fits in the 24GB of VRAM.
In most situations it's a better idea to just buy tokens. But there are definitely cases when that's not an option. And then a machine like the M5 Ultra can allow you to do things locally for a fairly limited budget. And in a simpler package to manage than a machine with multiple GPUs.
not getting on that bandwagon but wasn't that not the most demanding game as its a just a nonstop cutscene.
It's a very deliberate choice when if it doesn't make sense to gamers.
* it's critically acclaimed (86 metacritic, 10/10 IGN)
* whatever person decided this likely knows nothing about video games
* most importantly it's a modern game in UE5 that's COMING NATIVELY TO MAC, including to the App Store
What would you have put?
In US its $4000 upgade so $25 for 1GB.
Also:
> 512GB memory option for M5 Ultra coming late October
New Hampshire, Oregon, Montana, Alaska, Delaware.
So, on the mini the RAM upgrade runs at 25$ per GB on all tiers, the same as the Studio therefore the upgrade to 512 will probably cost 6400$.
The fully maxed out Apple Studio then will be 24699$. It's 17199 if you don't upgrade the storage.
Nevertheless I itch to have one :)
EDIT: or buy AAPL. If I had bought Apple stock instead of buying a Mac LC II in 1992, then I would have about $2 million in Apple stock.
That means for context under ~5k or so ttft (time to first token) it's going to respond faster than Claude. If the answer is less than ~1k I think the request finishes sooner. And it's ~claude 4.5 or 4.6 level intelligence.
Plus privacy. Plus offline.
For general inference there’s no ROI that makes this work vs subscriptions.
25k for computer now, plus 9-10% sales tax, plus operating cost, plus time and cost for R&D tinkering with models, harnesses, and infra (assuming highly capable engineering talent that can get paid for your human inference) vs a HEAVILY subsidized subscription at 200 per month with free R&D has a pretty long ROI (15 years?)
At API costs, it’s like 6 months if you’re heavy on inference. For training, specialized models will have their own ROI that makes this worthwhile. Then debate renting capacity and the platform to choose
If the pace of hardware change is what the jensens of the world say, then the rents on existing hardware will decline.
If you can settle for a M5 Max base model that would be a $49/month lease.
Today, you should be able to run Qwen 3.8 - 27B amazingly well on either which is giving comparable performance to 5.6 Luna on SWE Bench. The local models are now getting better and more efficient and this should give you headroom. Tools like turbo fieldfare are really reducing the memory requirements to run large models and I don’t see it stopping soon.
https://github.com/drumih/turbo-fieldfare
… I wish I hadn’t just calculated that.
Maybe they had the cash on hand, maybe they didn’t. Assuming they are in the family formation stage of life with young kids, an outlier risk appetite would be required to sleep well at night.
In RTP (NC), a ~$400k house at 5% down is $20k
You get what you pay for.
Depends on where you live. Also keep in mind population decline so long term, im not so sure. Prices already dropping in less desirable places (everywhere outside of most blue areas) (Assuming you are talking about the US sorry if not)
problem solved /s
Put together a similar build with a couple of rtx 6000 Ada cards and Apple's price tag suddenly looks pretty damn reasonable
They’re very different things.
The more logical argument to me is that Apple uses its upgrade price points as more than just direct BOM and rather as a proxy for things that are amortized across all their sales like support/warranty/etc so higher SKUs subsidize the costs of the lower ones.
And there are plenty of places in the US where you can buy a house for $150-$200k. Maybe not places you want to live for one reason or another, but they're there.
I wouldn’t recommend buying any bare metal unless money is a second thought or you can fully deduct the price.
Most often in the end you pay half the price then. Depending on the write offs you could even make some bucks out of it.
Or buy and lease. Under certain circumstances the hardware costs you nothing.
But you need money to save money. And a company.
to answer your question : looking at the aftermarket availability of Apple's prior best and brightest : practically no one buys them.
"people here buy them" , well, 'here' is one of the most affluent groups of people in the world.
They're available as movie and television set pieces (undoubtedly disappearing into the home of someone close to the staff post-production), and for administrative/boss types that can slip the cost into a ledger somewhere that few will ever see.
It has been a hobby of mine every few years to check out the apple site and see how big I can option a machine. My record was when I was in high school years ago and was able to option some pro studio-ish apple desktop thing to like 61,000 usd out the door.
It’s when self hosting and local hosting was the norm, and why it’s also starting to come back.
There will be workloads that can never touch a public cloud, and for it solutions like this are an option.
Even using multiple windows in parallel for as many as 5-10 hours per day, I find that I am not fully using my claude max (20x) and chatgpt pro (20x) accounts. I can for sure use up the claude max account, but chatgpt either gives me a free reset before I run out of tokens or I just fail to use the full quota. The quota for Sol seems like 10x that of Claude Opus at the same level, and forget Fable, you can use a 5 hour quota in 20 minutes.
But lets do the math:
Lets say a 20k workstation can run 1 inference at a time at the same speed you get with Sol hosted by openai (big assumption) and run an equally capable model (big assumption).
Each month this gives you about 100-170 inference hours on a Sol 20x Pro account, and 720 hours (if you utilize 24/7) on the workstation.
Assuming a 36 month amortization before the workstation has to be replaced due to no longer being able to run frontier models or is too inefficient due to electrical costs or what have you:
The monthly workstation cost is about $550 capex and $150 electricity -> $700/month
You would need about 6 Pro accounts to reach that capacity, which would cost you $1200 a month.
But this fails because: - You most likely can't utilize the workstation 24/7. Your work hours will be concentrated into 6-10 hours per day. - During work hours you are capable of utilizing more than 1 concurrent session. 6 Sol accounts would support as many as 20-30 during working hours, not all the time but if you could burst to that many (don't forget sub-agents and agent directed parallel agent workloads). - In 1 year the cost of Sol level models is likely to cost a fraction of what it does now.
this leads to:
The assumption previously used was that you can run a Sol level model on an M6 or whatever hardware $20k gives you. That is not true, it was an assumption made to show that even giving your own hardware every reasonable advantage it still loses.
Lets compare buying tokens of the best model you might run on your own hardware (still being unrealistic in favor of your own hardware) vs that same class of model on the market. I think one of the best models you might be able to run is GLM 5.4, but lets just look at chinese models generally:
$20k workstation, best case: $15k M5 Ultra 512GB, 36-month amortization, ~$440/mo. Runs a GLM-5.3-class model at ~30 tok/s. Saturated 24/7 it produces roughly 58M output tokens/month.
Buying those tokens:
So yes, at that speed for sure. But if the speed goes up? or the ability to batch at the same speed goes up? The economics start to shift. The gap is much closer, and you'd end up with a box you can still use or sell later.
Subscription pricing is still the best though!
Where you might win by owning your own hardware: - Hardware costs go up, and thus api costs go up. You've locked in your pricing. - Chinese/Open models become illegal/hard to access the way we do now. OpenAI and Anthropic are trying very hard to build a regulatory capture scheme to do this. I think they will be unsuccessful because China just won't participate.
For agentic coding, ~90% of the cost comes from cached input tokens. This cost increases quadratically with the session length. If sessions go near 1M context, the number of cached input tokens can easily exceed 1B in a day.
GLM-5.3 @ $0.26/M x 1000 = $260/day
This is the math to use.
isn't the whole point of all this ..... agents? isn't that what literally everyone is always clammering about in these threads? in which case the workstation is useful 720 hours out of 720 hours.
I have agents running 24/7 doing research, in fact I would argue this how they will be used for most programming tasks in the near future. For chatting, I agree local inference makes no sense. But for tasks that run continually, I'm not so sure. Personal computers took a while, local inference will too, but I think it will happen.
I'm not trying to cure cancer, although I do hope people who are use LLMs. ;)
That may not be many people, but there certainly will be some people who want to do that, and are willing to pay big bucks to do so.
If Apple didn't sold these things they wouldn't make them but, also the level of marketing that Apple is talking about for AI is basically the new group they need to capture because the ones I just listed are already buying Macs and or easily to motivate with the other obvious CPU / GPU performance upgrades for code compilation, faster memory and video transcoding.
AI based tools are very useful here - thinks like object removable or cleanup etc, not just AI generation.
For example Apple mentioned performance increases for https://learn.foundry.com/nuke/content/reference_guide/air_n...
Put another way: If $25k is the full extent of the start up capital costs, and operating costs are very low, that is a much cheaper business to start than most! The question is whether this is actually a useful model for a revenue generating business. I think that remains to be seen.
My guess is that there will be a few hits (which we'll hear a lot about - especially when someone actually pulls off "the first single-person unicorn", which I do suspect will happen someday) and a huuuge number of misses, which we won't hear much about.
This is, sadly, probably a foreign concept to a lot of people who have only worked at companies where hardware purchases are viewed as something to minimize and everyone is stuck with the same low spec laptops that the finance department picked out. At companies where someone might have a legitimate use for a $20K machine, their fully loaded costs (not their salary) are $300K or more, and other teams like sales are spending thousands of dollars per week on things like travel and hotels for their job, spending $20K on a computer that’s going to last several years is not a hard choice.
I'll find other ways to use the power though, opencode or maybe start working on more video projects.
I edit 4K ProRes and H.265 footage, sometimes with multicam (up to 4 streams) and color adjustments, titles, etc. It's only after stacking 3-5 effects before things can stutter, really.
Or if you try doing something CPU-intense in the background _while_ running some heavy creative software. I just don't do that.
I yield the floor to no one when it comes to pessimism, but that's incredible.
https://en.wikipedia.org/wiki/ELIZA_effect
It turns out the limiting factor isn't how sophisticated algorithms are, it's how gullible humans are.
Though we're pretty good at sizing up a person's emotional balance/maturity and competence at familiar tasks. So maybe have an old blacksmith watch the AI/robot interact with horse owners for a while, then shoe their horses, and see how well it does.
https://www.writingsbyraykurzweil.com/a-wager-on-the-turing-...
Those rules are fairly reasonable, at least as a start.
I've tried [1] and I almost 100% detect which is the AI. I really want to convince myself I have failed, does anyone know of a better site/resource for this?
I know it might be moving goalposts but I would consider AI to have passed in a well and truly undisputed manner when [2] is resolved.
But in a more practical sense, if AI can impersonate humans so well today then why are state of the art frontier models so obviously AI when they create PRs, commit messages, documentation, etc. Are the companies deliberately making them unnatural?
[1] https://turingtest.live/
[2] https://www.metaculus.com/questions/11861/date-when-ai-passe...
https://en.wikipedia.org/wiki/Turing_test
An the "popularized" version is faulty also since it uses an ideal, abstract human judge (like the "spheroidal economic agent").
But if you want to add declinations to the said popularized image of the Turing test, you may add Maxim Lott's IQ tests at trackingai.org . Around January 2025 LLMs reached an equivalent IQ of 100, for example.
I think there are elements showing lowering of performance and expectation.
Fab capacity is being bought online; there’s just lead time.
Noticeably greater intelligence is being achieved at the same number of parameters (see: Qwen3.8).
I think the future will be bright, it might be a matter of time. And for tinkers, a used Epyc + DDR4 server can be great fun and epic value.
If I sold just the two sticks of RAM in it right now, it’d pay for nearly half of the total cost.
I still have it, but never installed it because I only use ECC memory in my desktops. I'm saving it for a future Desktop in case memory prices never fall again.
However a problem exists with scalper bots. They're going to fight tooth and nail to keep prices artificially high even when supply increases. Same with the hyperscalers blowing "free money" at eating supply as a strategic position against smaller competitors.
[1] https://en.wikipedia.org/wiki/DRAM_industry_price_fixing
[2] https://moginlawllp.com/dram-antitrust-suit-ai-memory-supply...
Could you provide more details about the Epyc + DDR4 server?
The above is a standard project management problem. We do this for lots of industry all the time. There is every reason to think you can get a new factory running in 5 years.
Note that I said 1 factory above. Some of the special machines we don't have the ability to make them fast enough to do 2 (I don't know the real number!) new factories in 5 years. Existing factories are using most of the special machine capacity to replace machines that wore out on the way - this can be corrected as well, but it adds another year and the expenses are much larger. Realistically though 1 new factory is likely enough.
Basically everyone that makes memory is building new fabs, meanwhile I'm not sure how much longer AI datacenter demand for ram will last. I think the decrease in AI ram demand and the new fabs will likely coincide leading to a collapse in pricing.
That is, of course, assuming the memory manufacturers don't pull their favorite trick and collude.
https://transportgeography.org/contents/chapter3/transportat...
...
> I don’t think anyone truly knows when, but it will happen.
Do you know what cyclical means ... ?
Is there anything better now though?
All I see from AI, is an amplification of the enshittification of the internet.
And people being even more alone.
That's one big plus.
Compared to the prices of storage today, when people are presumably buying the product, it's actually a bad price.
You'll just have to be careful to match the enclosure to the ports on the system. The base-model M6 Mac mini still uses Thunderbolt 4, so a USB4v2 enclosure would be wasted.
You might be able to boot from the internal drive, install/upgrade the external drive, then boot from it.
Ok I'll be that guy. It's pretty easy to figure out if you're talking to an LLM now we know it's tics, failure modes, jailbreak techniques etc
Not being able to upgrade the SSD or the ram is a big issue, not so much on a portable laptop.
https://www.youtube.com/watch?v=IaVznUB6wss
The prices are ridiculous though. I may just keep rolling with my Windows 10 setup.
A m4 mac mini is better than al of these per dollar, msrp adjusted.
Hopefully by the end of the decade China figures out manufacturing at scale and fixes this.
Which one is it you can run local models on? I suppose the NPU only.
I can't believe that Apple still comes with this bullshit like 32 GBs is a lot. It's a lot for video memory - vRAM, but not RAM.
Is this a joke?
256 memory gets you to like 11k. So like 15-20k.
That's wild!
I wasn't able to debug network errors (restartin my Mac worked), Metal was missing low level disassembly / debugging tools (there is some hard to use UI), but the worst thing was the inflexible windowing system.
Even getting all the window handles on all screens/desktops with their titles and programs is impossible.
I just decided that I move to Omarchy 4 (basically Hyperland + QuickShell) + NVIDIA GPU, and I already was able to customize it more than my Mac in years.
I will miss Apple's hardware for sure, but not MacOS and the missing hardware documentation
It’s a bit depressing because it means that if I ever feel forced to switch my daily driver, it won’t come without a dump truck load of friction, frustration, and lost productivity, which I’ve validated by using the various Linux desktops on secondary machines.
It's still not well integrated of course as those plugins are from different people, but I at least don't feel powerless as I know I can make any change easily.
You can't expect any system, even the most configurable ones, to be ready out of the box to everyone's particular nitpicking without a minimum of effort.
https://pointieststick.com/2025/10/04/a-mac-like-experience-...
For example when using PyTorch I wanted to try to speed up my NN kernel by 2x by just using half precision and haven't noticed any speedup at all.
My main program missing from going back to Linux was ChatGPT Desktop, but now it's there.
I just checked out Hammerspoon, I'm happy for you that you wrote it, and looks great, but it has the same problem that I had: for security reasons Apple stopped allowing the window APIs to get all important information on other workspaces. You can only do it with Accessibility API. I was trying to fight with it but have up.
[0]: https://zen-browser.app/
The most disruptive thing in Omarchy is not the keyboard, it's the tiling WM in my opinion.
There's aren't really any alternative OS [1], but I don't think it's fair to say there's a closed software ecosystem. You can install/compile/run whatever you want, including kernel extensions. The hardware isn't locked.
[1] Hardware isn't locked: https://asahilinux.org/about/
> Apple allows booting unsigned/custom kernels on Apple Silicon Macs without a jailbreak! This isn’t a hack or an omission, but an actual feature that Apple built into these devices.
The performance you get with M5/M6 in theory is crazy (especially having matmul cores with such an energy envelope and unified memory). By the time I can use it well in Linux there will be M7/M8 though.
With NVIDIA I can buy a Blackwell based core right now and use it (even if I will have much less GPU memory. Though it will be high bandwidth).
They don’t provide any source code or documentation (which they easily could) so if want to do literally anything with your “open” device you must first reverse engineer the entire MacOS driver stack like the Asahi project did. Great fun, but it’s a criminal waste of talented human effort.
On top of that, the reason Asahi doesn’t work on M4 and later is that Apple has intentionally modified the ARM core to prevent stock MacOS from running once the hardware is “unlocked,” which makes it way more difficult to reverse engineer.
Bastards.
Naturally, the benefit for Linux is still the variety available in distributions. While IBM's RHEL and Canonical's Ubuntu may adopt some questionable stuff, Slackware, Devuan, Omarchy, and such likely won't.
Can you elaborate at all?
Between windows in a program is command+~
It is even more frustrating that this behavior is somewhat inconsistent, with some applications allowing this and others not (although it is more often the latter).
I made a workaround solution here with minimal setup/overhead: https://github.com/dainank/apple-click-through but I do wish someday this could be configurable in the OS as a setting.
The apps it works in are probably written in Cocoa née "Yellow Box" (allegedly so named because Bill Gates stated that rather than write apps for NeXT APIs he would instead....)
The way I fix it when using my Mac these days is to only use apps written using Objective-C and so forth.
Just amazing engineering push, the competition got the message and we benefit.