Design rack layout, network topology, and overall configuration for colocation
Design and implement an isolated Out-of-Band (OOB) management network (BMC, IPMI, Redfish) with full security hardening
Coordinate with colocation and hardware vendors — servers, switches, cross-connects, uplinks, remote hands, and cage/rack administration
Build bare-metal automation for fleet-wide zero-touch OS provisioning (MAAS / Tinkerbell / Cluster API)
Set up production-grade self-hosted Kubernetes clusters from scratch and maintain the core stack: Cilium (CNI), GitOps (ArgoCD), etc.
Own cluster lifecycle: upgrades, scaling, node maintenance, incident response
Write SOPs, runbooks, and post-mortems; participate in 24×7 on-call rotation
Establish hybrid-cloud interconnect between IDC and public cloud (VPN / Direct Connect / peering)
Pave the way for scale-out: GPU workloads (NVIDIA GPU Operator, passthrough, MIG) and VM workloads running alongside containers (KubeVirt or similar)
5+ years in data center / infrastructure / platform engineering
Hands-on experience building physical infrastructure from scratch: rack layout, network topology, server commissioning, and coordinating cross-connects and remote hands with colo / vendors
Practical experience designing and operating OOB management networks (BMC, IPMI, Redfish)
Have stood up production-grade self-hosted Kubernetes from scratch, and can independently debug cluster-level issues (CNI, CSI, storage)
Strong Linux systems administration and performance tuning (kernel, networking, storage I/O)
Bare-metal automation experience with at least one of: MAAS, Tinkerbell, Cluster API
Proficient with Terraform, Ansible, and at least one scripting language (Python / Go / Bash)
Experience with Cisco network and related techniques (VLAN, LACP/LAG, BGP, ACL, etc)
Experience with Palo Alto firewall configuration
Experience with storage systems (NetApp, Dell EMC, Pure Storage)
Fluent in English or Mandrain
High-density racks (30kW+) and 400G+ networking experience
Familiarity with immutable OS (Talos Linux / Flatcar / Bottlerocket)
Proficiency across both AWS and GCP; cross-cloud data migration experience
Experience building storage or Bigdata / offline data clusters
Virtualization experience (KubeVirt or similars)
Exposure to NVIDIA GPU Operator and K8s GPU workloads
CKA / CKS certification
CCNP / CCIE certification
Professional-level Japanese
World's largest crypto exchange by trading volume.
View company profileEstimated based on role seniority, stage (Private) & industry benchmarks.
You'll be redirected to the company's application page
Get roles like this daily
Join our Telegram channels for curated job alerts
Hey! Looking for your next role in Web3, AI, or Robotics? I can help.
Sign up to save jobs and access them across all your devices.