We are seeking a hands-on Senior AI Engineer who designs, builds, and operates production GenAI systems – agentic workflows, RAG pipelines, and LLM-backed services with real users and real SLAs. This is an engineering role, not a research role. The bar is reliability, latency, cost, observability, and safe deployment at scale, with end-to-end ownership from architecture through on-call. Typical workloads include enterprise knowledge platforms, conversational analytics, agentic automation, and LLM-augmented data products.
Experience the freedom of remote work from anywhere in Georgia, whether from the comfort of your home, our modern offices in Tbilisi and Batumi or a coworking space in Kutaisi.
Responsibilities
- Design agent orchestration (graph/state, conditional routing, tool calling, memory, checkpointing) in LangGraph / LangChain or equivalent
- Build production RAG end-to-end: chunking, embeddings, vector stores, hybrid retrieval, reranking, caching, and grounded synthesis
- Own Python / FastAPI services – async, SSE streaming, session handling, and structured error contracts
- Instrument with tracing and evaluation harnesses (MLflow, OpenTelemetry, or equivalent) for accuracy, cost, and regression
- Ship on Docker + Kubernetes (EKS/AKS/GKE) via CI/CD with test, eval, and canary gates
- Drive LLM cost engineering – model routing, prompt optimization, caching, token accounting, and build-vs-buy decisions
- Apply GenAI safety & governance: hallucination control, prompt-injection defense, PII handling, and HITL where required
- Partner with data engineering on semantic layers and pipelines (PySpark / SQL where applicable)
Requirements
- 5+ years in software engineering, with 2+ years shipping production LLM / agentic systems (not POCs or research)
- Proficiency in Python and FastAPI (async, REST, SSE)
- Production expertise in LangChain and LangGraph (or equivalent serious production experience with LlamaIndex, AutoGen, or MCP stacks)
- Background in production RAG: embeddings, chunking, and hybrid retrieval with reranking and caching
- Skills in vector databases such as Pinecone, Weaviate, pgvector, OpenSearch, or Databricks Vector Search
- Knowledge of at least one major LLM provider in production – AWS Bedrock (preferred), OpenAI / Azure OpenAI, or Anthropic – with model selection and routing trade-offs
- Competency in Kubernetes and Docker in real production environments (EKS/AKS/GKE)
- Expertise in cloud engineering on AWS
- Familiarity with observability and tracing tools (MLflow, LangSmith, OpenTelemetry), evaluation harnesses, and latency/cost ownership
- Capability to build CI/CD for AI systems (GitHub Actions, Jenkins, or equivalent) with test/eval gates
- Strong written and spoken English (B2 level); able to own design discussions with engineering and business stakeholders independently
Nice to have
- Databricks depth – MLflow (tracking & serving), Vector Search, Unity Catalog / Metric Views, PySpark / SQL
- Experience with LLM fine-tuning – PEFT, LoRA, QLoRA
- Understanding of MCP servers and tool integration
- Qualifications in GenAI governance & FinOps – auditability, prompt-injection hardening, PII, and token cost in regulated environments
- Background in classical ML / DL – NLP, BERT-family, time-series, and CV
We offer/Benefits
We connect like-minded people
- Delivering innovative solutions to industry leaders, making a global impact
- Enjoyable working environment, whether it is the vibrant office or the comfort of your own home
- Opportunity to work abroad for up to two months per year
- Relocation opportunities within our offices in 55+ countries
- Corporate and social events
We invest in your growth
- Leadership development, career advising, soft skills and well-being programs
- Certifications, including GCP, Azure and AWS
- Unlimited access to LinkedIn Learning and Udemy
- Free English classes with certified teachers
We cover it all
- Participation in the Employee Stock Purchase Plan
- Monetary bonuses for engaging in the referral program
- Comprehensive medical & family care package
- Five trust days per year (sick leave without a medical certificate)
- Benefits package (sports activities, a variety of stores and services)
EPAM Georgia is a team of innovators united by a passion for technology. The dynamic and inclusive culture we embrace helps positively impact our communities, clients, and employees. Here you will collaborate with multi-national teams, contribute to numerous cutting-edge projects, deliver the most creative solutions, and have an opportunity to learn. Our people are at the heart of our success, and we are proud to provide talents with a solid ground to develop and grow.
Why Choose Us
2024Best Place to Work 20242024Sitecore’s Partner Experience Awards
Read Full Description