HP ZGX Fury is now available to order, while HP, Red Hat and NVIDIA are planning a managed enterprise-AI stack that pairs the workstation with Red Hat AI Factory with NVIDIA. The September 9 announcement turns earlier hardware positioning into an edge-inference path, but sandbox timing, eligibility and supported configurations remain undisclosed.
- The ZGX Fury uses NVIDIA GB300 Grace Blackwell Ultra technology.
- HP quotes up to 20 PFLOPS FP4 and 748 GB unified memory.
- The hardware is orderable; the integrated Red Hat sandbox has no public date.
- Buyers should separate peak specifications from application-level throughput.
HP ZGX Fury availability and planned stack
HP’s official release says the workstation is certified for Red Hat Enterprise Linux and can be ordered now. The planned layer adds Red Hat AI Factory with NVIDIA so organisations can evaluate models and agents in a sandboxed environment before production.
HP ZGX Fury is the available hardware foundation; the integrated Red Hat AI Factory sandbox is a planned service whose timing and access details still need publication. That distinction prevents an orderable workstation from being mistaken for a fully available end-to-end platform.
| Element | Published fact | Buyer check |
|---|---|---|
| Compute | Up to 20 PFLOPS FP4 | Model-specific throughput |
| Memory | 748 GB unified | Usable capacity and precision |
| Software | RHEL certified | Support matrix |
| Sandbox | Planned | Date, eligibility and location |
Why HP is positioning local inference
HP frames the system for workloads that benefit from keeping compute close to users, machines and data. Manufacturing vision, engineering, branch operations and regulated environments may value lower network dependence or tighter data control. Local placement does not automatically lower cost or establish compliance; utilisation, support and integration decide the result.
The GB300-based system is designed for demanding inference and development. HP also describes workload isolation, scheduling and multi-GPU orchestration. Peak FP4 performance is useful for comparing a class of hardware, but teams should benchmark their actual model, context size, concurrency and accuracy requirements before treating the headline number as capacity.
Investing.com and MarketScreener separately carried the current collaboration and availability details. TechRadar’s earlier coverage provides context on the hardware class and enterprise pricing expectations, but the present HP release does not publish a price. Any procurement figure should come from an actual configuration quote.
Deployment questions for enterprise teams
Teams should verify supported models, quantisation, networking, storage encryption, high availability and how updates move from test to production. They should also ask whether the Red Hat layer is licensed, supported and monitored as one system or through separate contracts.
Governance follows the workload. A local coding assistant, vision model and autonomous operations agent need different permissions and evidence. Security teams should define what data leaves the device, which tools can be called and how model and configuration changes are audited.
The same system-level discipline appears in Arm’s agentic infrastructure launch and AWS unified routing control plane: architecture claims become useful only when operators attach measurements and rollback plans.
How to compare local and cloud inference
A useful comparison should hold the model, precision, prompt set and quality target constant. Teams can then measure first-token latency, sustained throughput, power use, administrator time and utilisation. A local system may reduce recurring inference charges, but idle capacity and specialist support can offset that advantage.
Data location also needs precise language. Running inference on the workstation can keep prompts and outputs on premises, yet software updates, licence checks, monitoring or connected retrieval services may still communicate externally. Buyers should map every data path and confirm which controls work when the system is disconnected.
Resilience testing should cover hardware failure, corrupted models and loss of the management plane. The business case is stronger when workloads have a documented recovery path and can move between local and central capacity without changing their security boundary. HP’s planned sandbox could help establish that path, but its final service design is not yet public.
Facilities planning is another practical constraint. Buyers should confirm rack placement, power, cooling, noise, network segmentation and physical access before delivery, then include those costs in the local-inference comparison. A workstation that meets a model benchmark can still miss its business objective if jobs queue behind one another or operational teams cannot maintain the environment. Pilot acceptance should therefore include sustained workload tests and a named support path, not only a successful demonstration.
FAQs
Is HP ZGX Fury available now?
HP says the workstation is available to order. It has not published access timing for the planned Red Hat AI Factory sandbox.
What chip powers HP ZGX Fury?
HP says it uses NVIDIA GB300 Grace Blackwell Ultra technology, with up to 20 PFLOPS FP4 performance and 748 GB unified memory.
Does local AI guarantee data privacy?
No. Local processing can reduce data movement, but access control, logging, updates and connected services still determine privacy and security.
Get the day’s top stories in your inbox
One concise email. No spam, unsubscribe anytime.



