AI/GPU Servers - 8x Custom VAST AI Servers
8 custom AI/GPU servers built for a client to rent out via Vast.ai. Supermicro TURIN2D24G, 2x EPYC 9754, 8x RTX 5090, 768GB DDR5 ECC.
About the Project
8 custom AI/GPU servers built for a client to rent out on the Vast.ai platform. Aura Digital is responsible for the full cycle: from hardware configuration design and component ordering, through assembly and stress testing, to data center deployment. Each server is optimized for AI training and inference workloads with maximum GPU density per rack unit.
Specification (per server)
- Chassis: Custom 6U rack-mount
- Motherboard: Supermicro TURIN2D24G-2L+
- Processors: 2x AMD EPYC 9754 (128 cores / 256 threads per server)
- GPU: 8x NVIDIA RTX 5090 (32GB GDDR7 per card)
- RAM: 12x 64GB DDR5 ECC (768GB total per server)
- Storage: 2x 4TB Samsung 990 PRO NVMe (8TB per server)
- Power: 5x CRPS 1600W (8,000W total, N+1 redundancy)
Project Scale
- 8 servers x 8 GPU = 64x RTX 5090 total
- 8 servers x 256 threads = 2,048 CPU threads total
- 8 servers x 768GB = 6,144GB RAM total
- Built for a client renting the capacity out via Vast.ai
Why This Configuration
A rental fleet is judged on two numbers: how many GPU-hours it can actually bill, and what it costs to keep those GPUs fed with work. Every decision below follows from that.
Why RTX 5090 and not a data-center card. For the inference and fine-tuning jobs that dominate marketplace demand, 32 GB of GDDR7 per card covers the majority of rented workloads at a fraction of the acquisition cost of an H100-class accelerator. Payback on consumer-class silicon is measurably shorter, and that is what the client is actually buying.
Why dual EPYC 9754 (256 threads). Eight GPUs starve if the host cannot feed them. Renters run data loaders, video decode and preprocessing on the CPU; a thin host becomes the bottleneck and the machine collects poor reliability scores. 256 threads and 12 memory channels per server keep the GPUs as the constraint, which is the whole point.
Why 768 GB of ECC RAM. Marketplace tenants routinely request large host memory for dataset caching. ECC is not negotiable: a silent bit-flip in a 30-day training run is an unbillable job and a support ticket.
Why N+1 power. 8,000 W across five CRPS units means a single PSU failure degrades the node instead of killing it. On a rental platform an unplanned reboot costs both the running job and the reliability score that drives future bookings.
Thermal and Density Engineering
Eight full-height cards in one chassis is not a mounting problem, it is an airflow problem. Cards sitting shoulder to shoulder ingest each other's exhaust, the middle GPUs throttle first, and the whole node's benchmark drops - visible to every prospective renter.
The build addresses this with a strict front-to-back pressure path: sealed cable routing so air cannot bypass the card stack, counter-rotating fan walls sized for the full 8,000 W thermal load, and validated inlet temperature at the middle slots rather than only at the chassis edge. Stress testing runs every GPU at sustained full load simultaneously - the realistic worst case on a rental platform - instead of one card at a time.
Delivery Process
- ROI model first. Before a single component is ordered we model achievable revenue against acquisition cost, power draw and colocation fees. If the configuration does not pay back, it does not get built.
- Component sourcing. EU supply, matched production batches, full warranty paperwork retained in the client's name.
- Assembly. Chassis preparation, PCIe topology check, cable management for airflow rather than for looks.
- Stress testing. All GPUs at full load simultaneously, thermal and power telemetry logged, memory validated.
- Deployment. Rack installation, network and remote management (IPMI) configuration, platform onboarding.
- Handover. Documentation, monitoring access, and a client who can operate the fleet without us.
Result
The same engineering approach later went into the NEOXIS server, which passed Vast.ai verification at 98.79% reliability - the platform metric that decides whether a machine gets rented at all. The methodology is documented in detail in our Vast.ai verification case study.
What Makes It Special
One of Aura Digital's largest hardware projects - 64 high-end GPU cards in a unified infrastructure. The process includes detailed ROI analysis for VAST AI profitability, component selection for the optimal price/performance ratio, and cooling design for uninterrupted 24/7 operation under full load.
Want the same for your own fleet? Configure a server or run the numbers in the GPU ROI calculator.
Have a similar project in mind?
Tell us - we'll scope it and return a quote within 2 business days.

