Skip to content
View MithunDataPro's full-sized avatar

Block or report MithunDataPro

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
MithunDataPro/README.md

Hi, I'm Mithun Dama

Senior AI/ML Engineer | Data Scientist | GenAI | Production ML Systems
πŸ“ Michigan, USA | Open to Hybrid Roles (Dearborn, Troy, Detroit)


About Me

Results-driven AI/ML Engineer and Data Scientist with 8+ years of experience building production-grade AI systems across automotive and enterprise domains.

Currently working at General Motors, designing and deploying scalable AI/ML and GenAI solutions using Python, cloud platforms (GCP/AWS), and modern architectures like RAG, LLMs, and MLOps pipelines to transform large-scale datasets into business-aligned decision systems.

Previously worked across data science and enterprise AI roles, solving real-world problems in domains like automotive, education, and industrial analytics.

Core Focus Areas:

  • Predictive Modeling & Advanced Statistical Analysis
  • Generative AI (LLMs, RAG, Multimodal Systems)
  • Gradient Boosting (XGBoost, LightGBM), Random Forests, GLMs
  • Feature Engineering & Model Optimization
  • Advanced SQL & Distributed Data Processing
  • Experimental Design & A/B Testing
  • MLOps & Production ML Deployment

Technical Expertise

πŸ”Ή Programming & Analytics

  • Python (Pandas, NumPy, Scikit-learn, SciPy)
  • Advanced SQL (multi-terabyte datasets)
  • PySpark
  • Statistical Modeling & Hypothesis Testing

πŸ”Ή Machine Learning & AI

  • Supervised & Unsupervised Learning
  • Gradient Boosting (XGBoost, LightGBM)
  • Random Forest
  • Logistic Regression & GLMs
  • Clustering & Dimensionality Reduction
  • Cross-Validation & Model Evaluation (MAE, RMSE, ROC-AUC)
  • Bias-Variance Optimization
  • Model Drift Monitoring

πŸ”Ή Generative AI & LLMs

  • RAG Architectures (Retrieval-Augmented Generation)
  • LLMs (Mistral, LLaMA, Multimodal Models)
  • Prompt Engineering
  • LoRA / QLoRA Fine-tuning
  • LangChain, LangGraph
  • Vector Databases (Qdrant)

πŸ”Ή Cloud & MLOps

  • GCP (BigQuery, Vertex AI)
  • AWS (S3, SageMaker, Glue, EC2)
  • Azure
  • Docker & Kubernetes
  • CI/CD (GitHub Actions, Jenkins)
  • Terraform
  • ML Lifecycle Management & Monitoring

πŸ”Ή Data Platforms

  • PostgreSQL, Snowflake, Cassandra
  • Hadoop Ecosystem (Spark, Hive)
  • Power BI & Tableau

Professional Experience

πŸš— General Motors | AI & ML Engineer

Feb 2023 – Present | Warren, MI

  • Architected and deployed GenAI and RAG-based systems for automotive engineering workflows, enabling intelligent data retrieval and reducing manual analysis effort.
  • Built multimodal AI pipelines integrating logs, documents, and system data using vector databases (Qdrant) and LLMs.
  • Designed scalable ML pipelines on GCP (BigQuery, Vertex AI) for processing large-scale automotive datasets.
  • Implemented LLM fine-tuning (LoRA/QLoRA) and optimization techniques (quantization, pruning) to reduce inference cost and improve performance.
  • Developed AI-powered applications such as log analyzers, document intelligence systems, and debugging assistants.
  • Operationalized models using Docker, Kubernetes, and CI/CD pipelines, ensuring production reliability and scalability.
  • Built monitoring and drift detection systems to maintain model performance in production.

πŸ“Š Yocket | Senior Data Scientist

May 2017 – Aug 2022 | India

  • Developed predictive models improving customer conversion by 30% using ensemble techniques (XGBoost, Random Forest) on large-scale CRM and financial datasets.
  • Built forecasting systems to estimate student demand and financial capability for study abroad planning.
  • Designed A/B testing frameworks improving campaign performance by 12% through statistically robust experimentation.
  • Engineered large-scale analytics pipelines using SQL and Snowflake to improve targeting precision by 15%.
  • Built ETL pipelines integrating multi-source data, reducing manual processing effort by 30%.
  • Delivered actionable insights to product and marketing teams, enabling data-driven business decisions.

🏭 Scon Design India Pvt Ltd | Senior Data Scientist

Aug 2018 – Jul 2021

  • Developed forecasting and predictive models for demand planning and process optimization.
  • Processed large-scale IoT and time-series datasets using Spark and Hadoop.
  • Implemented robust model validation using MAE, RMSE, and cross-validation.
  • Built distributed data systems handling multi-terabyte workloads.

Certifications

  • Databricks Certified Data Engineer Professional
  • Microsoft Certified: Azure Data Engineer Associate
  • AWS Certified Data Engineer – Associate

Areas of Interest

  • Scalable Predictive Modeling
  • Generative AI & Agentic Systems
  • Enterprise RAG Architectures
  • Cloud-Native ML Systems
  • Automotive AI & SDV Systems

Let's Connect

πŸ“§ mithundama.de@gmail.com
πŸ’Ό LinkedIn: https://www.linkedin.com/feed/
πŸ–₯️ GitHub: https://github.com/MithunDataPro


Professional Philosophy

Strong models don’t create impact β€” production-ready, validated, and scalable systems do.

Pinned Loading

  1. Real-Time-Streaming-with-Azure-Databricks-and-Event-Hubs Real-Time-Streaming-with-Azure-Databricks-and-Event-Hubs Public

    building real-time analytics streaming solution using Azure Databricks and Event Hubs.

    Jupyter Notebook

  2. Data-Engineer-Repo Data-Engineer-Repo Public

    Welcome to the Data Engineer Resources and Projects repository! Feel free to explore, use, and contribute to enhance the shared knowledge base for data engineers worldwide. Let's build and learn to…

    Jupyter Notebook 3

  3. Fraud-Analytics-using-Azure-Synapse-and-Power-BI-End-to-End-Project Fraud-Analytics-using-Azure-Synapse-and-Power-BI-End-to-End-Project Public

    In this end to end project , Fraud analytics is done using Synapse Analytics workspace and Power BI. ONXX regression model have been used to predict the fraud in Azure Synapse Sql dedicated pool.

    Jupyter Notebook

  4. Big-Data-Tools Big-Data-Tools Public

    To provide detailed information about each Big Data tool and concept

  5. Tokyo-Olympic-Azure-Data-Engineering-Project Tokyo-Olympic-Azure-Data-Engineering-Project Public

    The Dataset contains the details of over 11,000 athletes, with 47 disciplines, along with 743 Teams taking part in the 2021(2020) Tokyo Olympics. And Using ADF Extracting Data, and storing in ADLF …

    Jupyter Notebook 1

  6. End-to-End-Azure-Data-Engineering-Project End-to-End-Azure-Data-Engineering-Project Public

    In this project we are going to create an end to end data platform right from Data Ingestion, Data Transformation, Data Loading and Reporting.

    TSQL