ML LLM AI
We are looking for an ML Engineer to focus on developing and optimizing algorithms that accelerate large language model (LLM) inference. Your work will directly impact latency, cost efficiency, and scalability of production-grade AI systems. You will explore and implement cutting-edge techniques such as speculative decoding, prompt compression, quantization, and generation optimizatio
Company Sobolev Research Center
Sobolev Research Center is a local branch of an international full cycle IT company, based in Novosibirsk.
Must-have:
● Strong experience with deep learning frameworks (PyTorch or TensorFlow) ● Solid understanding of Transformer architectures and LLMs ● Experience with model inference optimization ● Strong Python skills ● Understanding of GPU/CPU performance and memory bottlenecks [Tech Stack] PyTorch, Hugging Face Transformers, TensorRT, ONNX Runtime, vLLM, SGLang, DeepSpeed, FlashAttention, xFormers, Quantization tools (BitsAndBytes, GPTQ)
Full time, office mode only, Novosibirsk, VMI
Log InOnly registered users can open employer contacts.
Our website uses cookies, including web analytics services. By using the website, you consent to the processing of personal data using cookies. You can find out more about the processing of personal data in the Privacy policy