Data Scientist

FourKites

Full Time3+ yearsPosted 18 days ago

Let the right jobs find you

Get personalised suggestions from verified company career pages, matched to your role, location, level, and skills.

Overview

Position Type

Full Time

Experience

3+ years

Job Description

What you'll be doing:

  • Design, build, and productionize ML models for problems like ETA/ATA prediction, using regression, classification, and time-series forecasting techniques
  • Develop NLP/LLM-based extraction pipelines for message-based ETA and status updates (text extraction, entity recognition)
  • Own models end-to-end: data pipeline → training → deployment → monitoring → retraining
  • Work with noisy, real-world logistics and supply chain data (GPS pings, check calls, carrier data) rather than clean, pre-processed datasets
  • Diagnose gaps between offline evaluation performance and live production accuracy, and drive fixes
  • Build and maintain automated training/retraining pipelines using orchestration tools such as Airflow
  • Set up and maintain model monitoring and observability (e.g., Grafana) to catch drift and degradation proactively
  • Replace manual or rule-based processes with ML-driven automation (e.g., automating manual check calls)
  • Translate model performance improvements into business impact — operational savings, efficiency gains, and deal-relevant outcomes
  • Mentor and guide other data scientists/engineers on technical approach and best practices
  • Make build-vs-buy and architecture tradeoff decisions independently

About the team:

Our product and engineering teams are dedicated to providing the industry's best-in-class end-to-end supply chain visibility platform. We are committed to building a high-performing, ML-driven team that turns supply chain data into automated, proactive action — and we want you to help lead that effort.

Who you are:

  • Strong ML fundamentals across regression, classification, and time-series forecasting
  • NLP experience — text extraction, entity recognition, or LLM-based extraction
  • Production ML experience — you've shipped models serving real traffic, not just built POCs or notebooks
  • Strong Python and SQL skills — pandas, scikit-learn, and comfort querying large datasets (Redshift/Snowflake a plus)
  • Experience with cloud and data infrastructure — AWS (S3, EC2), and orchestration tools like Airflow for training/retraining pipelines
  • Experience setting up or working with model monitoring and observability tooling (Grafana or similar)
  • Comfortable working with noisy, real-world data rather than clean, curated datasets
  • Experience diagnosing and closing the gap between offline evaluation results and live production performance
  • A track record of replacing manual/rule-based processes with ML solutions
  • Ability to translate model output into business value and communicate that impact to non-technical stakeholders
  • Experience collaborating cross-functionally with product, engineering, and operations teams
  • Experience mentoring or guiding other data scientists or engineers
  • Ability to make build-vs-buy and architecture tradeoffs independently
  • A track record of reducing manual intervention or turnaround time through automation
  • Excellent oral and written communication skills

Nice to have:

  • Experience in logistics, supply chain, or transportation
  • Familiarity with real-time/streaming data (Kafka)
  • Exposure to LLM/GenAI applications in production

Benefits:

  • Medical benefits start on first day of employment
  • 36 PTO days (Sick, Casual and Earned), 5 recharge days, 2 volunteer days
  • Home Office set ups and Technology reimbursement
  • Lifestyle & Family benefits
  • Mental Wellness support and guidance
  • Ongoing learning & development opportunities (Professional development program, Toast Master club, etc.)

Required Skills

Machine LearningRegressionClassificationsTime Series ForecastingNlpPythonSqlAwsAirflowGrafana

About the Company

FourKites

Chennai, India

Share This Job