I am a Data Scientist and Machine Learning Engineer focused on building rigorous, interpretable, and research-grade AI systems for biomedical signal processing, probabilistic modeling, and applied decision support.
My work combines software engineering, statistical learning, and scientific computing to design end-to-end pipelines: from data preprocessing and experimental validation to model selection, visualization, and reproducible reporting.
I specialize in machine learning for complex biomedical data, including EEG analysis, Hidden Markov Models, Gaussian Mixture Models, kernel methods, RKHS embeddings, and interpretable classification workflows. I also work with full-stack prototypes, SQL-based data pipelines, and AI-assisted product engineering.
Open To
- Teach Advance and Basic STEM Courses
- Data Scientist roles
- Machine Learning Engineer roles
- Applied AI / Biomedical AI research projects
- Research software engineering collaborations
- Interpretable ML and probabilistic modeling projects
| Domain | Proficiency | Details |
|---|---|---|
| Machine Learning | Advanced | Supervised and unsupervised learning, model evaluation, imbalanced data handling, cross-validation, classification pipelines |
| Probabilistic Modeling | Advanced | Hidden Markov Models, Gaussian Mixture Models, Markov processes, Bayesian reasoning, statistical inference |
| Kernel Methods | Advanced | RKHS embeddings, Gaussian kernels, precomputed kernels, similarity learning, distance-based classification |
| Biomedical Signal Processing | Advanced | EEG preprocessing, channel selection, filtering, z-score normalization, subject-level feature representation |
| Scientific Python | Advanced | NumPy, pandas, SciPy, scikit-learn, Matplotlib, Jupyter, Colab, experiment automation |
| Deep Learning | Intermediate / Advanced | PyTorch, TensorFlow/Keras, neural network workflows, representation learning, NLP foundations |
| NLP & LLMs | Intermediate / Advanced | Text preprocessing, tokenization, corpus analysis, LangChain, LLM-assisted workflows |
| Data Engineering | Intermediate / Advanced | SQL pipelines, PL/pgSQL routines, structured datasets, reproducible data preparation |
| Product Engineering | Intermediate | Full-stack prototypes, frontend interfaces, backend configuration, database integration, research-to-product workflows |
EEG-Based Supported Diagnosis of ADHD using Subject-Specific HMMs and Stationary RKHS Embeddings
A research-grade machine learning pipeline for classifying ADHD versus control subjects from EEG data using subject-specific generative modeling and RKHS-based similarity learning.
| Category | Details |
|---|---|
| Stack | Python, NumPy, SciPy, pandas, scikit-learn, Matplotlib, Optuna, Jupyter |
| Scale | 121 EEG subjects, frontal EEG channels, HMM-GMM topologies N3G3, N4G4, N5G5 |
| Performance | Nested CV balanced accuracy approximately 73.5% with 95% CI and permutation testing |
| Security | No raw clinical data redistributed; reproducible notebooks and controlled dataset access |
| Impact | Interpretable biomedical AI workflow for EEG-based ADHD support |
| Repository | leonlpz/EEG-Based-Supported-Diagnosis-of-ADHD-using-Subject-Specific-HMMs-and-Stationary-RKHS-Embeddings |
This project models each subject with a dedicated Hidden Markov Model with Gaussian Mixture emissions, then compares subjects through closed-form RKHS distances and Probability Product Kernels. The resulting similarity matrices are used by KNN and SVM classifiers with precomputed kernels, enabling a statistically rigorous and interpretable classification workflow.
Mastering NLP from Foundations to LLMs
A practical NLP and LLM learning repository focused on modern language-processing workflows, classical NLP foundations, text classification, embeddings, and large language model applications.
| Category | Details |
|---|---|
| Stack | Python, pandas, Matplotlib, NLP pipelines, LLM tooling |
| Scale | Multi-chapter NLP codebase covering foundations through LLM systems |
| Performance | Educational and experimental repository for reproducible NLP workflows |
| Security | Public learning repository; no sensitive data embedded |
| Impact | Supports development of NLP, LLM, and applied AI engineering skills |
| Repository | leonlpz/Mastering-NLP-from-Foundations-to-LLMs |
This repository strengthens the engineering foundation required to build NLP applications, including preprocessing pipelines, text classification, embeddings, mathematical foundations, and LLM-oriented workflows.
PyTorch NLP Book
A deep learning and NLP-focused repository for practical experimentation with PyTorch-based language models, text pipelines, and neural network implementations.
| Category | Details |
|---|---|
| Stack | Python, PyTorch, NLP, Jupyter |
| Scale | Notebook-driven NLP and deep learning examples |
| Performance | Practical implementation repository for model experimentation |
| Security | Public educational codebase |
| Impact | Reinforces applied deep learning and NLP engineering practice |
| Repository | leonlpz/PyTorchNLPBook |
This repository supports experimentation with neural NLP systems and helps bridge statistical learning, deep learning, and practical AI implementation.
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow
A machine learning practice repository focused on production-relevant ML concepts, classical learning workflows, neural networks, and end-to-end experimentation.
| Category | Details |
|---|---|
| Stack | Python, scikit-learn, Keras, TensorFlow, Jupyter |
| Scale | Comprehensive ML learning and implementation repository |
| Performance | Experiment-oriented machine learning workflows |
| Security | Public learning repository |
| Impact | Builds applied ML engineering capability across classical and deep learning models |
| Repository | leonlpz/Hands-on-Machine-Learning-with-Scikit-Learn-Keras-and-TensorFlow |
This repository consolidates applied ML workflows, from feature engineering and model evaluation to neural-network experimentation and reproducible notebooks.
Aug 2024 — Present
Development and validation of a multimodal machine learning system integrating neurophysiological and biomedical data with interpretable representations to identify patterns associated with impulsivity-related mental disorders.
Scope of Work
- Built data pipelines for psychometric, EEG, and biomedical datasets
- Developed probabilistic models using PCA, GMM, HMM-GMM, RKHS distances, and kernel methods
- Implemented classification workflows using KNN and SVM with precomputed kernels
- Designed visualization routines for model interpretation and experimental analysis
- Supported reproducible research documentation, methodological reporting, and scientific validation
Python SQL Machine Learning EEG HMM-GMM RKHS scikit-learn Biomedical AI
2024
Instructional role focused on software engineering foundations, programming logic, applied computing, and technical mentoring.
Scope of Work
- Supported students in software development fundamentals
- Guided practical implementation of programming concepts
- Connected mathematical reasoning with computational problem solving
- Reinforced clean coding, documentation, and reproducible workflows
Teaching Software Engineering Python Programming Logic Technical Mentoring
2021 — 2024
Academic teaching experience in mathematics, physics, and technical STEM training, with emphasis on analytical reasoning, modeling, and bilingual instruction.
Scope of Work
- Delivered mathematics and physics instruction in bilingual academic environments
- Designed didactic material for analytical and quantitative reasoning
- Supported student learning in trigonometry, algebra, calculus, and scientific thinking
- Integrated computational and applied examples into classroom explanations
Mathematics Physics Bilingual Education STEM Curriculum Design
| Recognition | Details |
|---|---|
| CONACYT Scholarship | Full-time graduate scholarship for M.Sc. studies at CINVESTAV Unidad Monterrey |
| Biomedical AI Research | Developed interpretable EEG-based ADHD classification pipeline using HMM-GMM and RKHS embeddings |
| ACEMATE Research Program | Contributor to national research initiative on impulsivity-related mental disorders |
| SPIE Student Member | Participation in scientific and engineering community activities |
| Big Data Technical Training | 527-hour Big Data technical diploma from Fundación Carlos Slim |
| Scientific Computing Portfolio | Public GitHub portfolio focused on ML, probabilistic modeling, NLP, and biomedical signal processing |
Learning:
- Advanced machine learning systems
- Probabilistic graphical models
- NLP and LLM engineering
- Biomedical signal processing
- Reproducible research software
Building:
- EEG-based interpretable AI pipelines
- HMM-GMM and RKHS classification workflows
- Multimodal machine learning systems
- Scientific computing notebooks
- Research-ready documentation
Exploring:
- Kernel methods for distribution comparison
- Generative modeling for subject-specific biomedical data
- LLM-assisted scientific workflows
- Applied AI for healthcare and decision support
Open To:
- Teach Advance and Basic STEM Courses
- Data Scientist roles
- Machine Learning Engineer roles
- Biomedical AI collaborations
- Research software engineering projects