As an NLP / LLM Engineer at Rubiscape, you
will build the language intelligence that makes RubiAI — our agentic AI copilot
— genuinely useful to enterprise analysts, data engineers, and business
decision-makers interacting with complex data in natural language. You will own
the full NLP stack: from classical text processing and information extraction
through LLM fine-tuning, retrieval-augmented generation, and prompt engineering
for enterprise-grade accuracy and safety. This role directly advances Rubiscape’s
vision of a platform where every employee can query, analyse, and act on
enterprise data without writing a single line of code.
· Design and
implement NLP pipelines for document understanding, entity extraction,
classification, and summarisation integrated into RubiAI and Business
Workbench.
· Fine-tune
and evaluate open-source LLMs (Llama 3, Mistral, Phi-3) on Rubiscape’s
domain-specific corpora using PEFT techniques (LoRA, QLoRA) and managed
fine-tuning APIs.
· Build and
iterate on RAG architectures: chunking strategies, embedding model selection
and evaluation, hybrid retrieval (dense + sparse), and re-ranking for
high-precision enterprise knowledge retrieval.
· Develop
and maintain the prompt engineering framework for RubiAI: structured prompting
patterns, chain-of-thought templates, system prompt governance, and automated
regression testing for prompt quality.
· Evaluate
LLM outputs systematically using RAGAS, DeepEval, or custom evaluation
harnesses; track quality metrics across model versions and prompt iterations.
· Implement
multilingual NLP capabilities for Indian-language enterprise content (Hindi,
Marathi) where required by government and public-sector customers.
· Collaborate
with the AI Platform Engineer to deploy NLP microservices at scale, ensuring
latency SLOs are met for interactive RubiAI copilot responses.
· Experience
with structured output generation (JSON mode, grammar-constrained decoding) for
reliable LLM integration with enterprise data systems.
· Knowledge
of speech-to-text or multimodal models relevant to voice-driven analytics
interfaces.
· Exposure
to LLM safety, alignment, and responsible AI practices for regulated-sector
deployments.
· Contribution
to open-source NLP or LLM tooling (HuggingFace, LangChain, LlamaIndex, or
equivalent).
Rubiscape is India’s leading Decision
Intelligence Platform, unifying data engineering, BI, machine learning, and
agentic AI in a single governed platform. Built in Pune and trusted by Fortune
500 enterprises across BFSI, manufacturing, healthcare, and government. 8
international innovation patents. 10 Industry-Academia Labs & COEs. From BI
to AI — One Platform. Every Decision.
RequirementsRequirements
· 3+ years
of NLP engineering experience with demonstrable production deployments; at
least 1 year working with large language models in a product context.
· Deep
proficiency with HuggingFace Transformers, Datasets, and PEFT libraries for
model fine-tuning and evaluation.
· Hands-on
experience with RAG pipeline construction using LangChain or LlamaIndex and at
least one vector store (Weaviate, Qdrant, Chroma, or pgvector).
· Strong
Python skills and familiarity with model serving for NLP workloads (ONNX,
TorchServe, vLLM, or Text Generation Inference).
· Solid
understanding of LLM evaluation methodology: benchmark construction, human
preference alignment, and hallucination mitigation strategies.
· Bachelor’s
or Master’s degree in Computer Science, Computational Linguistics, or a related
field.


