Mathematics in Data Science / Machine Learning

Mathematical modeling for reliable, interpretable data science.

I combine a strong mathematical foundation with applied machine learning, statistical evaluation, and careful communication. My work focuses on turning complex data into models and analyses that are technically sound, transparent, and useful in practice.

M.Sc.
Mathematics in Data Science
1.2
B.Sc. Data Science GPA
BMW
Data science internship
Machine Learning Statistics Deep Learning Model Interpretability
Portrait of Jan Stüwe
Jan Stüwe / Munich
Profile Data Scientist
Focus
Statistics, probability, interpretable machine learning
Current
M.Sc. Mathematics in Data Science at TUM
Approach
Mathematical rigor, empirical validation, clear reporting
01 Statistical inference
02 Model evaluation
03 Scientific computing
Statistical modeling Machine learning Probability theory Scientific computing Research communication

Work experience

Data Science Intern
02/2025 to 04/2025
BMW AG, Munich
  • Built clustering pipelines for automotive data, partitioning, hierarchical, density based.
  • Evaluated quality with internal and external metrics, reported findings.
  • Applied stratified sampling for representativeness and stable comparisons.
  • Ran descriptive and inferential analysis to validate conclusions.
Python Clustering Statistics Reporting
Student Assistant
10/2023 to 04/2025
Catholic University of Eichstätt-Ingolstadt, Ingolstadt
  • Graded assignments for Foundations of Machine Learning (Chair of Reliable Machine Learning, Apr 2025 – Aug 2025).
  • Graded assignments for Analysis I and Analysis II (Chair of Mathematical Analysis, Sep 2023 – Mar 2025).
  • Translated and LaTeX-typeset course scripts for Stochastics and Scientific Computing (Chair of Data Assimilation, Sep 2024).
  • Translated and LaTeX-typeset the Analysis 3/Integration Theory course script (Chair of Reliable Machine Learning, Jul 2023).
LaTeX Grading

Projects

Lorenz 63 data assimilation project visualization
Machine learning

Neural network assisted 3DVar on Lorenz 63

In this project, we implemented a neural-network-supported 3DVar data assimilation pipeline on the chaotic Lorenz ’63 system. We generated a synthetic twin experiment and computed innovations and 3DVar increments to establish a strong classical baseline. On top of that, we trained a compact MLP (ReLU, Adam, early stopping) to predict analysis increments directly from the background state and observations, and also evaluated a hybrid scheme that blends the learned increment with the 3DVar update. Evaluation included an 80/20 train–validation split, Monte-Carlo studies across multiple observation-noise levels, and a partial-observations setting. Across settings, the learned model consistently reduced mean L2/RMSE versus standard 3DVar. Overall, the approach demonstrated robust improvements in state estimation over classical 3DVar.

PyTorch Data assimilation Evaluation
Car sales prediction project visualization
Applied ML

Time-to-sell prediction for cars

As part of a team of four, I worked on a project for Audi focused on predicting the time it takes to sell cars. We developed a deep learning model that accurately predicted the timespan, achieving a margin of error of just 9 days, significantly outperforming the baseline of approximately 20 days. To ensure the predictions were interpretable, we also utilized a Generalized Additive Model (GAM), which allowed us to gain valuable insights into the key factors influencing the sales times. The project placed a strong emphasis on handling high-dimensional data and performing comprehensive feature preprocessing to ensure optimal model performance. Our team presented our findings multiple times at Audi, demonstrating the model’s effectiveness and impact. The project was graded with a top score of 1.0.

Deep learning GAM Feature engineering
Fake news classification project visualization
NLP

Fake News Classification

In my second semester, I worked on a project focused on classifying news as fake or true of this Kaggle Fake and Real News Dataset, utilizing both classical machine learning classification methods and modern deep learning approaches. The project required extensive data preprocessing, feature engineering, and careful model evaluation to ensure accuracy and robustness. By applying a combination of traditional techniques and advanced neural networks, I was able to achieve highly reliable results. The final model achieved an accuracy of approximately 99%, demonstrating strong generalization across both classes. The project was graded with the highest possible score of 1.0.

NLP Transformers Validation

Education

Master of Science, Mathematics in Data Science

Technical University of Munich, Munich

10/2025 to present

  • Focus: Probability theory
  • Topics: advanced math, modeling, Python

Bachelor of Science, Data Science

Catholic University Eichstätt-Ingolstadt, Eichstätt

10/2022 to 09/2025

  • Specialization: applied math, scientific computing
  • Tools: Python, R, MATLAB, Git
  • GPA: 1.2

Skills

Core
Python MATLAB R NumPy pandas Matplotlib Scikit-Learn
ML
PyTorch TensorFlow Keras XGBoost Interpretability
Engineering
Git GitHub Linux Docker pytest SQL MySQL
Web
HTML CSS
Cloud
AWS
Writing
LaTeX
Python
Python
MATLAB
MATLAB
R
R
NumPy
NumPy
pandas
pandas
Matplotlib
Matplotlib
Scikit-Learn
Scikit-Learn
PyTorch
PyTorch
TensorFlow
TensorFlow
Keras
Keras
Git
Git
GitHub
GitHub
Linux
Linux
MySQL
MySQL
HTML
HTML
CSS
CSS
AWS
AWS
LaTeX
LaTeX

Want the full story behind a project?

Send me a message!

Contact

Don't hesitate to reach out!

Location
Munich, Germany
Links