As a Data Scientist/Machine Learning Engineer, you'll finetune models, improve data quality, add new signals, and push solutions to production.
Sumble is building a knowledge graph from web data with a first focus on data for go-to-market teams. We use sources like job posts and resume data to identify things like org structure, tech stack, and key projects (e.g., GenAI initiatives, cloud migrations). Our product already has strong product-market fit, early revenue, and happy customers — and now we’re ready to accelerate.
Our long-term vision is to become the primary destination for accessing high-quality web data. Try the product at sumble.com.
Our Team: We are a team of 15, including 10 engineers with experience at companies such as Google, Meta, Stack Overflow, and Kaggle.
What you'll do- Finetuning small language models
- Improving the quality of existing data using scalable approaches. Examples include: making sure URLs are associated the right company, we have the correct HQ address, we have mapped parents-subsidiary using techniques like LLM validation, SERP, and triangulating across sources.
- Adding new signals: this usually involves scrubbing, matching and normalizing new signals and matching to our existing ontology
- Pushing solutions into production environments, which may involve touching data pipelines and/or backend systems
Requirements
- Located within Americas timezones
Our Tech Stack:
- ML/Data: PyTorch, Huggingface, Gemma models, LORA, VLLM, Skypilot, Marimo
- Languages & Frameworks: Python, FastAPI, React, Typescript
- Cloud Platform: Google Cloud Platform (GCP)
- Databases: PostgreSQL, DuckDB
- Infrastructure: Cloud Run
- Product/Design: Figma, Vercel V0
Challenges We Tackle:
- Transforming noisy datasets into high-quality data products
- Running expensive analytics computations efficiently
- Managing the complexity of a growing number of data sources, machine learning models, and large data operations
- Create a great PLG experience with upsell pathways
Benefits
- Medical, dental, and vision (US)
- 401k (US)
- Target 4 weeks PTO
Similar Jobs
Software
Product-oriented senior ML role to build, deploy, and operate models and ML systems (classification, extraction, entity resolution, clustering, ranking, anomaly detection, forecasting, LLM pipelines). Define datasets, labeling, evaluations, monitoring, and collaborate with engineering and product to ship production-quality training/inference pipelines and model-serving infrastructure.
Top Skills:
Apache BeamSparkAWSAzureClaude CodeGCPLlmsNumpyOpenai CodexPandasPythonPyTorchRagScikit-LearnSQLTensorFlowTool-Augmented Agents
Artificial Intelligence • Information Technology • Software
As a founding Data Scientist/Machine Learning Engineer, you'll develop AI/ML models, enhance product capability, and drive impactful user outcomes while working closely with product teams.
Top Skills:
Data ScienceMachine Learning
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Develop and maintain project schedules and resource plans; coordinate tasks, milestones, and meetings; monitor progress, budgets, and risks; produce status reports; implement process improvements; ensure projects meet scope, timeline, and quality requirements.
Top Skills:
AgileDcma 14-PointExcelMS OfficeMs ProjectSpmWaterfall
What you need to know about the Pune Tech Scene
Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.



