Designs and optimizes production-grade Retrieval-Augmented Generation systems for enterprise AI applications. Responsibilities include architecting retrieval pipelines, optimizing context windows and token usage, implementing re-ranking and semantic caching, building document chunking and ingestion workflows, governing hybrid vector search, and measuring retrieval accuracy, hallucination rates, latency, and cost. The role requires deep expertise in Python, vector databases, embeddings, LLM orchestration, SQL, and cloud data engineering, plus a mandatory professional cloud or analytics certification.
This is a remote position.
RAG Architect
Job Details
- Employment Type: Contract
- Work Mode: Remote
- Location: Offshore
- Total Experience Required: 6 to 10 years
- Relevant Experience Required: 3+ years of dedicated experience designing production-grade Retrieval-Augmented Generation (RAG) architectures and optimizing LLM token throughput
- Mandatory Certification: Google Cloud Certified Professional Cloud Database Engineer, AWS Certified Data Analytics - Specialty, or Databricks Certified Data Engineer Professional
Job Summary
We are seeking an experienced Context Window Optimization / RAG Architect to take full ownership of our enterprise generative AI retrieval performance, accuracy, and operational cost metrics. The ideal candidate will design high-throughput knowledge retrieval systems, optimize semantic context parsing, build custom re-ranking pipelines, and engineer caching grids to deliver data-grounded AI responses with minimal latency and maximum token efficiency.
Key Responsibilities
- Architect end-to-end advanced Retrieval-Augmented Generation (RAG) pipelines, building structures for document parsing, semantic metadata enrichment, and multi-vector lookups.
- Optimize context window utilization patterns, designing smart parent-child chunking models, sentence-window retrievals, and sliding window strategies to eliminate irrelevant text tokens.
- Build high-performance re-ranking layers, deploying machine learning cross-encoders (e.g., Cohere Rerank, BGE-Reranker) to score retrieved documents before feeding them into the LLM context pool.
- Implement automated semantic caching architectures, utilizing caching layers (e.g., GPTCache) to capture recurring semantic queries, reducing API token expenditures and response latencies.
- Establish automated data chunking pipelines, configuring ingestion routines to cleanly parse semi-structured and unstructured formats (PDFs, corporate wikis, SQL outputs) into clean vector targets.
- Govern vector similarity spaces, fine-tuning hybrid search algorithms that cleanly combine dense semantic embeddings with sparse keyword token indexes (BM25).
- Audit context-level hallucination rates and accuracy logs, tracking precision metrics, retrieval recall bounds, and processing speeds to systematically eliminate incorrect model generations.
Requirements
- 6 to 10 years of enterprise data engineering, database design, or search engine engineering experience, with 3+ dedicated years actively scaling context retrieval loops for live LLM applications.
- Strong technical mastery of Python, vector databases (Pinecone, Milvus, Weaviate), text embedding models, open-source orchestration tools (LlamaIndex, LangChain), and SQL.
- Deep structural understanding of context window limitations ("lost in the middle" phenomenon), multi-modal token dynamics, network data transfer speeds, and cloud memory spaces.
- Mandatory certification: Professional Cloud Data/Database Engineer or Specialty Analytics credential from a major cloud vendor (AWS/GCP/Azure).
Preferred Qualifications
- Prior experience implementing Graph RAG frameworks utilizing native knowledge graphs (e.g., Neo4j) to map complex corporate data relationship networks.
- Familiarity with fine-tuning open-source text embedding models specifically optimized for industry-specific terminology or legacy product schemas.
Similar Jobs
Artificial Intelligence • Information Technology • Machine Learning • Software • Virtual Reality • Analytics
Leads enterprise AI and Generative AI solution architecture across pre-sales, proposals, RFPs, discovery workshops, demos, prototypes, and proof-of-concepts. Designs scalable LLM, RAG, agentic AI, cloud, data, and integration architectures; supports estimation, feasibility, and risk analysis. Develops reusable AI accelerators, reference architectures, thought leadership content, and innovation roadmaps while mentoring pre-sales teams and engaging executive stakeholders.
Top Skills:
Agentic AiApi IntegrationsAws Ai/MlAzure Ai ServicesAzure OpenaiData PipelinesGcp AiLarge Language Models (Llms)MicrosoftMlopsPrompt EngineeringRetrieval-Augmented Generation (Rag)Servicenow
Artificial Intelligence • Hardware • Information Technology • Machine Learning
Develops and controls integrated project schedules for large-scale EPC, industrial construction, semiconductor fab, infrastructure, and related projects. Responsibilities include baseline management, critical path and delay analysis, schedule variance monitoring, recovery planning, contractor schedule validation, progress forecasting, stakeholder workshops, and executive reporting. The role coordinates engineering, procurement, construction, and commissioning teams while using Primavera P6, project controls methods, BI dashboards, and AI tools to support delivery decisions.
Top Skills:
ChatgptClaudeMicrosoft CopilotExcelMicrosoft PowerpointMicrosoft ProjectPower BIPrimavera P6Tableau
Artificial Intelligence • Hardware • Information Technology • Machine Learning
Manages indirect material planning, procurement coordination, inventory accuracy, tooling orders and maintenance, and WIP management for semiconductor assembly and test operations. Drives manufacturing improvements through AI, automation, dashboards, and real-time issue resolution. Maintains process documentation, supports quality and regulatory compliance, and collaborates cross-functionally to improve production flow and achieve operational KPIs.
Top Skills:
Artificial IntelligenceAutomationDashboardsData Visualization SoftwareMesSAP
What you need to know about the Pune Tech Scene
Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.


.jpeg)