LiveRemoteFull-timeApply by 1 Nov 2026
Data Engineer (LLM, RAG)
Remote
- Experience
- 5–15 years
- Employment
- Full-time
- Work mode
- Remote
- Salary
- Not disclosed
- Deadline
- Apply by 1 Nov 2026
- Posted
- 2025-04-15
Required skills
| Skill | Experience | Level |
|---|---|---|
| LLM API | 5+ years | Expert |
| LLM integration | 5+ years | Expert |
| Large Language Model (LLM) | 5+ years | Expert |
| RAG | 5+ years | Expert |
| Python | 5+ years | Not specified |
| Apache-Kafka | 5+ years | Expert |
| Apache Spark | 5+ years | Expert |
| Hadoop | 5+ years | Expert |
| Machine Learning | 5+ years | Expert |
| NumPy | 5+ years | Expert |
| Pandas | 5+ years | Expert |
| Pytorch | 5+ years | Expert |
| SciKit-Learn | 5+ years | Expert |
About the role
JOB OVERVIEW:
Job Title: Data Engineer (LLM, RAG)
Company: Trantor Inc
Experience: 5-10 years
Location: Remote Working
No. of Positions: 1
MANDATORY CRITERIA:
- Strong Data Engineer Profile
- 5+ Years of Experience in statistical machine learning, deep learning, data mining
- 2+ Years of Experience working with large language models (LLMs)
- Proficiency in Python + at least one other programming language
- Prompt engineering and retrieval-augmented generation (RAG) techniques
- Experience with frameworks and libraries like PyTorch, Numpy, Pandas, SciPy, Scikit-Learn, LangChain, and Hugging Face Transformers
JOB RESPONSIBILITIES:
- Design, train, fine-tune, and deploy LLMs using prompt engineering and RAG techniques
- Build scalable solutions using frameworks and libraries like PyTorch, Hugging Face Transformers, and LangChain
- Collaborate with cross-functional teams to deliver data-driven solutions
- Develop efficient data pipelines using big data technologies like Spark and Hadoop
- Ensure compliance with data privacy and ethical AI practices
REQUIRED SKILLS:
- Python programming with working knowledge of Java, Scala, or C++
- Experience with LLMs, including training and deployment
- Prompt engineering and RAG techniques
- Frameworks and libraries like PyTorch, Numpy, Pandas, SciPy, Scikit-Learn, LangChain, and Hugging Face Transformers
- SQL and big data technologies like Hadoop and Spark
- Cloud platforms like AWS, Azure, or Google Cloud
IDEAL CANDIDATE:
- Master’s or Ph.D. in Computer Science, Statistics, or related field
- 5+ years of experience in machine learning, deep learning, and statistical modeling
- Strong proficiency in Python and other programming languages
- Expertise in PyTorch, Hugging Face, LangChain, Scikit-Learn, Pandas, NumPy, and SciPy
- Strong analytical and problem-solving abilities
- Effective communication skills and ability to work in cross-functional teams
WHAT TO EXPECT:
- Remote working with flexible work arrangements
- Opportunity to work with a large-scale/global company