I have one. The limitation (beyond total size of the memory) with the Spark is DDR5. "Real" inference hardware is HBM (high bandwidth memory) which is like 10x the performance.
So for prefill -- which is more about compute than bandwidth -- the Spark performs quite admirably. But on decode it's highly bandwidth constrained. Some smaller MoE models (e.g. Gemma4) can do 60-70 tok/second but anything dense, or anything that is actually filling up most of that 128GB is going to choke out around 15 tok/sec. Even at NVFP4.
For my current work I get to log into trays on a real GB300. It's somewhat comical that NVIDIA is marketing the little baby on my shelf here as even in the same universe as that. Which is basically like having access to a super computer.
Not everyone is handling PII. Where I work, anything like that is only available to a very limited set of people who absolutely need to be able to see it. Also some systems allow access control at the column and even row level, so even if it's intermingled with other data you want the LLM to read, you might be able to mask it that way.
Also, people shouldn't be running any LLM on data of a business without a proper contract in place like you have with any vendor who has access to your data. And if there's specific PII requirements, those should be covered too.
A system-agnostic language for magic spells with a compiler capable of producing the magic wand movements, incantation, hand signs or magic circle required to perform the spell
the trouble with compact is that no one really knows how it works and what it does. hence, for me at least, there is just no way I would ever allow my context to get there. you should seriously reconsider ever using compact (I mean this literally) - the quality of CC at that point is order of magnitute significantly worse that you are doing yourself significant disservice
if you actually hit the compact (you should never be there no matter what but for the sake of argument) more often than not you'll see CC going off the rails immediately after compacting is done. it even doesn't know what it did let alone you :)