Staff Software Engineer, AutoCloud, Context and Memory
Summary
Staff engineer acting as the principal technical authority for Google Cloud's AutoCloud portfolio, designing scalable agent memory architectures, context caching, hybrid search/RAG platforms, and distributed low-latency systems powering AI-driven autonomous cloud operations.
AutoCloud is Google Cloud’s autonomous, AI-powered cloud management portfolio. We are transforming how enterprise customers design, deploy, operate, investigate, and optimize their workloads and infrastructure across Google Cloud Platform (GCP). In this role, you will be the principal technical authority guiding the design of scalable memory architectures, solving complex state retrieval issues, and partnering with Principal Engineers, researchers across DeepMind, and partner teams across Google Cloud to deliver a high-precision, low-latency, and secure context platform for autonomous operations.
Google Cloud accelerates every organization’s ability to digitally transform its business and industry. We deliver enterprise-grade solutions that leverage Google’s cutting-edge technology, and tools that help developers build more sustainably. Customers in more than 200 countries and territories turn to Google Cloud as their trusted partner to enable growth and solve their most critical business problems.Individual pay is determined by factors including job-related skills, experience, and relevant education or training.US: $207000 - $300000 (USD) + 20% bonus target + equity + benefits
Learn more about benefits at Google.
- Own the end-to-end architecture, technical roadmap, and core goal for agent memory systems, dynamic context synthesis pipelines, graph-based cloud representations, and hybrid search/RAG platforms.
- Lead the design and implementation of low-latency context caching, token compression/pruning strategies, working memory buffers, and long-term episodic knowledge stores for autonomous agents.
- Architect high-throughput, low-latency distributed systems and automated benchmarking frameworks to ensure sub-second cloud state aggregation, high retrieval recall, and hallucination mitigation.
- Ensure all context and memory subsystems meet stringent enterprise-grade multi-tenancy standards, tenant data isolation policies, compliance mandates, and fine-grained access controls.
- Partner across research (e.g., DeepMind) and platform service teams to standardize shared context models and APIs, while mentoring engineers and upholding architectural review standards.
Minimum qualifications:
- Bachelor's degree in Computer Science, AI/ML, Data Systems, Information Retrieval, a related technical field, or equivalent practical experience.
- 8 years of experience in software development.
- 5 years of experience testing, and launching software products, and 3 years of experience with software design and architecture.
- 5 years of experience leading ML design and optimizing ML infrastructure (e.g., model deployment, model evaluation, data processing, debugging, fine tuning).
- 2 years of experience with GenAI techniques (e.g., LLMs, Multi-Modal, Large Vision Models) or with GenAI-related concepts (language modeling, computer vision).
- 2 years of experience building infrastructure on cloud platforms.
Preferred qualifications:
- Master’s degree or PhD in Engineering, Computer Science, or a related technical field.
- Experience building low-latency, high-availability distributed storage systems and APIs on major cloud platforms.
- Expertise in agent memory (working/episodic), context caching, token pruning, vector search, and knowledge graphs.
- Ability to define technical roadmaps, author comprehensive design docs, and align multi-organization stakeholders. Demonstrated skill in coaching engineers and clearly communicating complex architectures to leadership and research partners.
- Track record in hybrid search, graph databases, and querying complex cloud telemetry and topology.
- Background in engineering secure, multi-tenant cloud architectures with strict data isolation and compliance controls.