Hardware Engineer
Apply NowAbout aionAion is the enterprise AI platform, a full-stack solution for building, fine-tuning, and deploying AI at scale. Whether an organization is modernizing internal operations, launching AI-powered products, or transforming customer experiences, Aion takes them from concept to production on a single, unified platform. We work differently than most AI companies: our teams deploy alongside our customers, turning production-ready AI into real business outcomes in weeks, not quarters. We’re a fast-growing, VC-backed startup led by founders with a track record of successful exits. With teams across the US, UK, and India, we’re building the next generation of enterprise AI and we’re looking for exceptional people to help us scale. Who You AreYou're a hardware engineer passionate about building the infrastructure that powers large-scale AI systems. You understand the complexities of modern compute platforms including servers, GPUs, networking, storage, and high-performance computing environments that enable enterprise AI workloads. You have experience designing, integrating, validating, and optimizing hardware platforms for performance, scalability, and reliability. You're comfortable working across compute infrastructure, hardware bring-up, system diagnostics, firmware interactions, and data center deployments. You're product-minded and understand how hardware decisions directly impact AI performance, infrastructure efficiency, and customer experience. You enjoy solving complex systems challenges and collaborating with software, platform, and infrastructure teams to build production-ready AI infrastructure. What You'll DoHardware Platform Design & Integration • Design and build hardware platforms optimized for AI training and inference workloads. • Evaluate, integrate, and validate servers, GPUs, networking equipment, storage systems, and accelerator hardware. • Develop scalable hardware architectures supporting enterprise AI deployments across cloud, hybrid, and on-premises environments. • Collaborate with software and platform engineering teams to optimize hardware and software integration.
System Performance & Optimization • Optimize compute, memory, storage, networking, and GPU performance for AI workloads. • Benchmark and analyze system performance under production-scale AI deployments. • Identify hardware bottlenecks and implement improvements to maximize throughput, latency, and infrastructure efficiency. • Validate hardware compatibility across multiple AI frameworks and deployment environments.
Infrastructure Reliability • Build reliable and fault-tolerant hardware infrastructure capable of supporting mission-critical AI workloads. • Develop hardware validation, diagnostics, stress testing, and failure analysis processes. • Support hardware lifecycle management including provisioning, upgrades, maintenance, and replacement strategies. • Collaborate with vendors and partners to evaluate emerging hardware technologies.
Deployment & Operations • Support hardware deployment across enterprise customer environments and internal infrastructure. • Develop automation and operational procedures for hardware provisioning, monitoring, and troubleshooting. • Ensure infrastructure meets enterprise standards for availability, security, and operational excellence. • Work closely with customer-facing teams to resolve deployment and infrastructure challenges.
Engineering Excellence • Establish best practices for hardware validation, documentation, testing, and operational readiness. • Conduct technical reviews and contribute to infrastructure architecture decisions. • Collaborate across hardware, software, platform, and AI engineering teams to continuously improve system performance and reliability.