Linux Infrastructure Engineer
Uvation
Job Overview
We are seeking a highly experienced Senior Linux Infrastructure Engineer with deep expertise in Linux administration, bare metal infrastructure, enterprise storage, and next-generation AI Factory / GPU infrastructure platforms . This role is focused on designing, deploying, operating, and troubleshooting large-scale Linux-based infrastructure that powers both traditional enterprise workloads and modern AI/ML environments.
This is not a DevOps-focused role . We already have a dedicated DevOps team and are looking for an engineer with extensive hands-on experience in Bare Metal as a Service (BMaaS), GPU infrastructure, high-performance storage, data center operations, and enterprise Linux platforms .
The ideal candidate will have experience building and managing infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI-ready platforms. They should be comfortable working with high-performance computing (HPC), AI Factory environments, and large-scale Linux deployments where performance, reliability, and operational excellence are critical.
Key Responsibilities & Required Skills
Linux & Bare Metal Infrastructure
- Expert-level Linux administration (Ubuntu required; Red Hat and SUSE preferred)
- Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management
- Experience operating Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments
- Strong understanding of server hardware, including:
- BIOS/UEFI
- RAID controllers
- Firmware management
- iLO/iDRAC/IPMI
- NICs and SmartNICs
- HBA cards
- Hardware diagnostics and troubleshooting
- Experience designing, implementing, and supporting enterprise Linux infrastructure at scale
AI Factory & GPU Infrastructure
- Experience deploying and managing GPU-accelerated infrastructure for AI/ML workloads
- Understanding of NVIDIA GPU technologies including:
- A100, H100, H200, B200, or equivalent GPU platforms
- NVIDIA DGX and OEM GPU servers
- GPU provisioning and lifecycle management
- GPU monitoring and performance optimization
- Knowledge of AI Factory architecture and infrastructure requirements
- Experience supporting GPU clusters, AI training environments, and high-performance computing (HPC) workloads
- Understanding of:
- GPU resource allocation and scheduling
- Multi-GPU systems
- GPU networking requirements
- High-bandwidth, low-latency infrastructure design
- Familiarity with NVIDIA ecosystem technologies such as:
- CUDA
- NCCL
- GPUDirect Storage
- NVIDIA Fabric Manager
- NVIDIA Base Command (preferred)
Enterprise Storage & Data Platforms
- Advanced Linux storage administration:
- LVM
- XFS, EXT4
- NFS
- iSCSI
- Fibre Channel SAN
- Multipath I/O
- Strong hands-on experience with Ceph , including:
- Cluster architecture
- MON, OSD, MDS
- RBD, CephFS, RGW
- Capacity planning
- Performance tuning
- Failure recovery
- Experience with high-performance AI storage platforms such as:
- WEKA
- VAST Data
- Dell PowerScale
- Pure Storage FlashBlade
- NetApp
- Understanding of:
- NVMe-over-Fabrics (NVMe-oF)
- RDMA
- GPUDirect Storage
- Parallel file systems
- AI data pipelines
Networking & Infrastructure
- Strong networking knowledge:
- Bonding
- VLANs
- Routing
- MTU optimization
- DNS
- DHCP
- Experience with high-performance data center networking:
- 100G/200G/400G Ethernet
- RoCE
- RDMA
- Spine-Leaf architectures
- Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or equivalent technologies
- Strong understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting
Operations & Reliability
- Experience with high availability, clustering, and disaster recovery
- Strong troubleshooting skills across:
- Linux operating systems
- Hardware platforms
- GPU infrastructure
- Networking
- Enterprise storage
- Experience supporting mission-critical production environments
- Bash and Python scripting for automation and operational efficiency
- Experience creating operational documentation, runbooks, and infrastructure standards
- Understanding of AI infrastructure design and reference architectures
- AI cloud integration for workloads
- SOP and runbook development and maintenance
- Incident, problem, and capacity management
- Business continuity and disaster recovery planning for AI workloads
- Proactive risk identification and mitigation to avoid business impact
Nice to Have
- Kubernetes infrastructure (especially AI/ML and GPU integration)
- KVM, VMware, OpenShift Virtualization, or similar virtualization platforms
- Ansible automation
- NVIDIA Base Command Manager
- Slurm or HPC workload schedulers
- Observability and monitoring platforms (Prometheus, Grafana, OpenTelemetry)
- Data Center Infrastructure Management (DCIM) tools
- IPAM solutions
- AWS, Azure, or hybrid cloud exposure
We Are Not Looking For
- Candidates whose experience is primarily CI/CD pipeline engineering
- Engineers focused mainly on Terraform, GitOps, or application delivery pipelines
- Cloud-only administrators with limited bare metal, storage, or hardware experience
- Professionals whose primary expertise is software development rather than infrastructure engineering
Ideal Candidate
Someone who has spent years designing, building, and operating enterprise Linux environments, large-scale bare metal infrastructure, storage platforms, and modern AI Factory environments. The ideal candidate understands how to deploy and manage GPU-enabled infrastructure, BMaaS platforms, enterprise storage, and high-performance networking while solving complex operating system, hardware, storage, and AI infrastructure challenges. DevOps experience is a plus, but deep Linux, infrastructure, storage, BMaaS, and AI Factory expertise is the primary requirement.
- ...We are seeking a skilled and experienced Azure DevOps Engineer with a strong background in Linux administration to join our dynamic team. The ideal... ...responsible for managing and optimizing our Azure cloud infrastructure, ensuring seamless CI/CD pipelines, and maintaining...SugeridoTiempo completo
- ...others. We are looking for a Backend Engineer to join our Core team and help build... ...of production systems. Contribute to infrastructure improvements, including CI/CD, observability... .... ~ Experience working with Linux and Docker. ~ Ability to write and maintain...SugeridoTemporalTrabajar en la oficinaRemotoTrabajo híbridoHorario flexible
- ...build what makes people better and keep challenging ourselves to inspire others. Your impact: Manage and optimize cloud infrastructure and resources with attention to cost and resilience. Enhance technical monitoring and alerting to ensure the availability...SugeridoTiempo completoTrabajar en la oficinaRemotoTrabajo híbridoHorario flexible
- ...Job Overview We are looking for an experienced Network Security Engineer to design, implement, monitor, and support enterprise security infrastructure across on-premises, cloud, and hybrid environments. The ideal candidate should possess strong expertise in next-generation...SugeridoTiempo completoRemotoTrabajo híbrido
- ...communication. Preferred: experience optimizing UI rendering and memory management and building applications for macOS, Windows, Linux, or other platforms. Bonus: contributions to open-source projects, especially in P2P or decentralized technology. Benefits ~...SugeridoTiempo completoRemoto
- ...Service, BI, Tagging & Tracking, Data Ops, and the nova product/engineering teams. Key Responsibilities: - Design, build, and maintain... ...tools, to accelerate development and build intelligent data infrastructure. Document what works so it becomes a standard team pattern....Tiempo completo
- ...billion in equity and debt funding from global and regional investors, and is now valued at $6,5 billion. We are hiring a Data Engineer for the Data Services Team who will help us create our internal solutions such as a Feature Store and will be ready to perform...Tiempo completoContratoRemoto
- ...We are seeking a Senior iOS Engineer to own the client side implementation of how members join, pay, and stay with Raya. Member Experience owns the full member lifecycle – applying, onboarding, payments, and lifecycle management – and this role owns the surfaces where...Tiempo completo
- ...standards like the Model Context Protocol (MCP). As a Backend Engineer, you will be instrumental in executing rapid prototyping and... .... AI Observability & MLOps: Familiarity with cloud-native infrastructure (AWS/GCP/Azure), containerization (Docker/Kubernetes), and AI...Trabajo remotoContratistaTiempo completoContrato
- ...Position: Full Stack Engineer (AI Integration) Location: Remote from LATAM Contract Type: Full-time Contractor Time Zone Alignment: CT About In All Media We are a Managed Nearshore Teams provider headquartered in Austin, specializing in building and...Trabajo remotoContratistaTiempo completoContrato
- ...Position: Senior Fullstack Engineer Location: Remote from LATAM Contract Type: Full-time Contractor Time Zone Alignment: PST / PDT (Pacific Time) About In All Media We are a Managed Nearshore Teams provider headquartered in Austin, specializing in building...Trabajo remotoContratistaTiempo completoContrato
- ...Position: Senior PHP & WordPress Engineer Location: Remote from LATAM Contract Type: Full-time Contractor Time Zone Alignment: CT About In All Media: We are a Managed Nearshore Teams provider headquartered in Austin, specializing in building and embedding...Trabajo remotoContratistaTiempo completoContrato
- ...project involves designing complex schemas, optimizing queries, and deploying automated pipelines to support advanced data science and engineering efforts. As a critical member of the team, you will drive production support, implement agile continuous delivery methodologies,...Trabajo remotoContratistaTiempo completoContrato
- ...Position: Senior Data Engineer, Cloud Economics Location: Remote from LATAM Contract Type: Full-time vendor (via Time Zone... ...the Cloud Economics Team to focus on engineering scalable data infrastructure capable of processing high-volume telemetry and...Trabajo remotoTiempo completoContrato
- ...high-impact digital products. Our teams specialize in software engineering, data platforms, artificial intelligence solutions, and... ...Overview We are seeking a Senior UI/UX Developer (Front-End Engineer) to drive the development of highly interactive, responsive, and...Trabajo remotoTiempo completo
- ...Position: Backend Engineer (AI Agentic) Location: Remote from LATAM Contract Type: Full-time Contractor Time Zone Alignment: CT About In All Media We are a Managed Nearshore Teams provider headquartered in Austin, specializing in building and embedding...Trabajo remotoContratistaTiempo completoContrato
- ...Platforms Team Our Content Platforms Team powers critical web infrastructure and user-facing content delivery across a high-scale global... ...Content Platform Design (UCMS): Collaborate with engineering and content operations teams to design, optimize, and maintain...Trabajo remotoTiempo completoEmpleo permanenteContrato
- ...releases, implementing robust version control strategies, and leveraging advanced DevOps platforms like Copado. As a Salesforce Release Engineer, you will be instrumental in bridging the gap between development, QA, and operations, ensuring that new features and metadata are...Trabajo remotoContratistaTiempo completoContrato
1 000 USD por ano
...notice and remotely. Support and maintain an AlwaysOn availability group (2 servers: primary + mirror). Help analysts and ML engineers with data access and retrieval. Run routine database maintenance: archiving, recovery strategies, security policies, data...Tiempo completoDesde casaRemoto- ..., research artifacts, and actionable stakeholder readouts. Cross-Functional Alignment: Partner closely with Product, Design, Engineering, Operations, and Service Management teams to validate tooling decisions and optimize workflows. Research Artifact Traceability...Trabajo remotoTiempo completoContrato
- ...opportunities. Cross-Functional Collaboration & Traceability: Partner with Product Managers, Solution Architects, Developers, and QA engineers to ensure requirements are understood, maintaining end-to-end traceability from business goal to technical delivery. UAT...Trabajo remotoTiempo completoContrato
- ...Position: Full-stack Engineer Location: Remote from LATAM Contract Type: Full-time (Vendor) Time Zone Alignment: ET About In All Media We are a Managed Nearshore Teams provider headquartered in Austin, specializing in building and embedding high-performing...Trabajo remotoTiempo completoContrato
- ...Functional Bridge & Stakeholder Management: Facilitate discussions between business stakeholders, Salesforce Administrators, and data engineering teams to align technical delivery with business objectives. Greenfield Platform Implementation: Contribute directly to a "...Trabajo remotoTiempo completoContrato
- ...the end-to-end recruitment lifecycle for complex IT roles. Your challenge will be to identify, vet, and present high-level software engineers and architects to US-based stakeholders, ensuring a seamless cultural and technical fit within modern, fast-paced environments....Trabajo remotoTiempo completoContrato
- ...system components and live production releases. Cross-Functional Collaboration: Partner with Product Managers, UX Researchers, and Engineers to validate concepts using qualitative/quantitative data, align on project milestones, and perform post-build Design QA. Must-...Trabajo remotoPrácticaTiempo completoContrato
¿Desea recibir más vacantes?
Suscríbase y reciba vacantes similares a Linux Infrastructure Engineer. ¡Sea el primero en aplicar!
