ML engineer (LLM quantization & optimization)

RUB 220,000450,000/month
Office
Full-time

ML LLM AI

Brief description of the vacancy

We are looking for an ML Engineer to focus on developing and optimizing algorithms that accelerate large language model (LLM) inference. Your work will directly impact latency, cost efficiency, and scalability of production-grade AI systems.  You will explore and implement cutting-edge techniques such as speculative decoding, prompt compression, quantization, and generation optimizatio

About the company

Company Sobolev Research Center

Sobolev Research Center is a local branch of an international full cycle IT company, based in Novosibirsk.

Responsibilities

  1. LLM Quantization & Low-Precision Optimization (AWQ, GPTQ, SmoothQuant, BitsAndBytes, INT8/INT4/FP8) Develop and apply quantization techniques to reduce model memory footprint and accelerate inference while preserving output quality. What you will do: Implement and evaluate weight-only and weight-activation quantization methods, including AWQ, GPTQ, SmoothQuant, and related approaches. Analyze how different quantization algorithms work and compare their accuracy, latency, memory usage, calibration requirements, and hardware efficiency. Work with symmetric and asymmetric quantization, as well as per-tensor, per-channel, and group-wise quantization schemes. Understand and apply both post-training quantization (PTQ) and quantization-aware training (QAT), selecting the appropriate approach for different models and deployment scenarios.
  2. Speculative Decoding & Generation Acceleration (EAGLE3, dFlash, dSpark, dFly, MTP) Design algorithms that reduce the number of decoding steps and improve generation speed. What you will do: a. Implement speculative decoding pipelines (draft + target models) b. Develop multi-token prediction approaches c. Explore parallel and tree-based decoding strategies
  3. Prompt Compression & Context Optimization (Token pruning / attention-based filtering, semantic compression via embeddings, LLM-based summarization (self-compression)) Reduce input context length without degrading output quality. What you will do: a. Compress long prompts and conversation history b. Filter irrelevant tokens dynamically c. Optimize context window usage

Requirements

Must-have:

● Strong experience with deep learning frameworks (PyTorch or TensorFlow) ● Solid understanding of Transformer architectures and LLMs ● Experience with model inference optimization ● Strong Python skills ● Understanding of GPU/CPU performance and memory bottlenecks [Tech Stack] PyTorch, Hugging Face Transformers, TensorRT, ONNX Runtime, vLLM, SGLang, DeepSpeed, FlashAttention, xFormers, Quantization tools (BitsAndBytes, GPTQ)

Working conditions

Full time, office mode only, Novosibirsk, VMI

Contacts

Log InOnly registered users can open employer contacts.

Our website uses cookies, including web analytics services. By using the website, you consent to the processing of personal data using cookies. You can find out more about the processing of personal data in the Privacy policy