Senior LLM/RAG Engineer
Summary
Senior/Principal technical consultant designing LLM and RAG architectures for enterprise clients. Work spans RAG pipelines, audio/video processing, multi-model tiered routing, LLM FinOps, and hands-on benchmarking across Chinese (DeepSeek, Qwen, Doubao, GLM, Kimi) and Western frontier LLM providers.
This is a PRE-SALE ACTIVITY. Preferably to be arranged free of charge
Position name: Technical Consultant — LLM/RAG Architecture
Level: Senior / Principal / Architect
Hard skills requirements (including years):
Production experience designing and running RAG pipelines: chunking, embeddings, vector databases, retrieval, reranking, evaluation
Audio/video processing pipelines: transcription (ASR), speaker diarization, timestamp indexing, handling multi-participant content
Multi-model / tiered routing architecture — routing lightweight tasks to cheap high-throughput models and escalating complex reasoning tasks to flagship models
Hands-on evaluation and benchmarking of LLM providers, including Chinese providers (DeepSeek, Alibaba Qwen, ByteDance Doubao, Zhipu GLM, Moonshot Kimi) alongside Western frontier models
LLM FinOps: token-level cost modelling, unit economics per processing run, and cost projection across a growth curve (10K → 175K → 400K users)
Capacity and throughput planning: rate limits, batching, caching, concurrency at enterprise volume
Awareness of the risks of Chinese LLM providers: data residency, compliance, latency outside China, vendor lock-in