科技職缺
職缺列表
Senior ML Engineer (Token Factory)
Senior ML engineer building fine-tuning and inference for foundation models on a large-scale GPU cloud, working in Python/JAX on LoRA/full-parameter training, speculative decoding, MoE architectures, and low-precision (FP8/NVFP4/MXFP4) methods across multi-node GPU systems.
Senior Network Engineer
Senior Network Engineer designing and operating large-scale data center, PoP, and backbone networks for an AI cloud platform, including InfiniBand GPU cluster interconnects and network automation tooling.
Senior Site Reliability Engineer — Token Factory (Inference Platform)
Senior SRE on Nebius Token Factory, owning reliability, performance, and observability for a GPU-accelerated AI inference platform running tens of thousands of GPUs. Core stack: Kubernetes, Prometheus, Grafana, Terraform, Python/Bash, with GPU serving tools (vLLM, Triton, Ray).
Senior Software Engineer (Token Factory)
Senior engineer building Nebius's Token Factory — a large-scale LLM training and inference platform within their AI cloud. Designs distributed systems in Python and Go, maintains ML infrastructure, and improves job scheduling across high-load web services.
Senior Technical Program Manager - New Data Center Launches
Senior TPM owning end-to-end delivery of new data center and GPU cluster launches at hyperscale, coordinating engineering, networking, construction, and operations teams across AI cloud infrastructure deployments in Europe.
Senior Technical Project Manager - Token Factory
Senior Technical Project Manager for Nebius's Token Factory, coordinating cross-functional AI infrastructure projects—new region launches, GPU capacity expansion, and end-to-end onboarding of new AI models—across engineering, SRE, infrastructure, networking, and security teams.
Senior Technical Project Manager - VPC
Senior Technical Program Manager driving complex cross-team initiatives for Nebius AI Cloud's Virtual Private Cloud (VPC) team — covering networking services like routing, security groups, load balancing, and VPN — across new region launches, hardware generations, and platform scalability.
Service Delivery Manager
Service Delivery Manager acting as the technical and operational interface between hyperscale customers and Nebius infrastructure teams, ensuring SLA compliance and reliable delivery of GPU cloud and bare metal AI infrastructure in Nebius data centers.
Site Reliability Engineer in Hardware Infrastructure
SRE on Nebius's Hardware Automation team building internal platforms and tooling for large-scale data center infrastructure. Day-to-day work centers on ensuring fault-tolerance, scaling services, implementing CI/CD, and troubleshooting hardware, software, and networking issues using Linux, Python, and Bash.
Site Reliability Engineer (SRE) AI Infrastructure (Early Career)
Early-career SRE assisting with day-to-day network infrastructure operations, deploying approved changes, and executing small SRE projects at an AI cloud company. Core stack includes Python/Go/C++, Linux, Kubernetes, Terraform, and low-level networking (eBPF, DPDK, TCP/IP).
Site Selection & Colocation Manager – Data Centers
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through…