I work at the intersection of Data Science, Artificial Intelligence, software engineering, automation, and cloud-oriented systems. I am especially interested in building solutions that are not only technically solid, but also useful in real-world environments.
My current focus is on machine learning, predictive modeling, ETL and data pipelines, applied software development, and MLOps/DevOps-oriented workflows. I care about systems that are maintainable, reproducible, and designed with practical impact in mind.
I am also expanding my engineering stack through cloud-based workflows and tooling, with hands-on exposure to Google Cloud and AWS, and a growing interest in learning Go.
- Machine Learning & Predictive Modeling
- Data Analysis & Statistical Reasoning
- ETL, Data Pipelines & Automation
- MLOps / DevOps Foundations
- Cloud-Oriented Systems
- Applied AI Systems
- Software for Real-World Decision Support
Best-ranked system at EXIST 2026 Task 2 (CLEF 2026), with three first places on the official leaderboards — 1st of 144 in binary sexism detection, and 1st of 117 and 1st of 186 in source-intention classification, topping Spanish and English simultaneously.
The core idea is to use Gemini as a semantic mediator rather than a black-box classifier. A meme is not one signal but two that only mean something together — an image and its text, where the sexism often lives in the tension between them. So each meme is interpreted offline into structured text (description, sexism analysis, reasoning, intention and irony cues, zero-shot probabilities), and that text is concatenated with the OCR, encoded with XLM-RoBERTa or a multilingual Longformer, and fused with EEG and Ekman emotion features before a probabilistic decision.
That single choice — letting the LLM translate rather than judge — lifted Task 2.1 validation AUC from 0.741 with a raw ViT to 0.880, the largest gain anywhere in the pipeline. Two other decisions carried the result: training on the empirical distribution of the six annotators instead of majority labels (Learning with Disagreement), justified by a mixed-effects analysis showing annotator demographics significantly shape sexism perception; and a fixed 0.6·model + 0.4·Gemini probability blend that produced the #1 runs outright.
The write-up is also honest about what failed: Task 2.3 ranks 10th of 118 on soft evaluation but 132nd of 187 on hard, and the paper diagnoses the gap as validation-overfit category thresholds and unmodeled label co-occurrence — not the model.
📄 Paper (CLEF 2026 Working Notes)
Python · Gemini · XLM-RoBERTa · Longformer · multimodal fusion · soft-label training
Classifies 31,634 scientific abstracts against the nine Planetary Boundaries — the Earth-system processes that bound a safe operating space for humanity, six of which have already been transgressed. Institutions increasingly have to evidence their environmental contribution, and "we publish on sustainability" is not evidence. The useful question is which Earth-system processes does this research actually study? Graded 9.9/10.
That question is harder than it looks, and the whole project is built around why: a paper can be about sustainability without studying a boundary, and can advance one without ever using SDG language. Words like climate or water routinely appear as background or motivation, not as the thing being measured. Keyword matching over-fires; a naive LLM is worse, happily labelling anything that sounds green. Every design decision is a trade between positivity bias (inflating the university's apparent contribution) and false negatives (understating real strengths).
Three findings make it worth reading. TF-IDF beats every transformer (F1 0.61 vs SPECTER's 0.51) — because abstracts state PB terminology explicitly and the PB reference documents are themselves keyword lists, so TF-IDF compares two representations of the same nature. Model choice went to Qwen 2.5 14B, not the most accurate one: Llama 3.1 8B had higher agreement but 0.00% rigorousness — it never rejected irrelevant documents, disqualifying for institutional mapping. And prompt engineering beat model size: the v4 "operational object" principle — classify by the variable the abstract actually measures, not by its narrative framing — raised Top-1 from 63.3% to 72.1% without returning to earlier versions' permissiveness, correcting 15 documents against v3's 4 in a paired comparison.
The auditable agent cascade is reported as it performed: it does not beat the single model, trading ~1 point of Top-1 for traces a human can review. Ships as a containerized platform (FastAPI + Next.js + Nginx, docker compose up, no GPU) with precomputed embeddings and an optional local RAG chatbot.
Python · Qwen 2.5 / Ollama · SPECTER2 · TF-IDF · agent cascade · FastAPI · Next.js · Docker
Decision-support tool for Valencia City Council, and a full MLOps cycle end to end — graded 10/10. 24 CatBoost models (one per hour of the day) predict traffic across 1,158 measurement zones, while an independent integer-programming optimizer (PuLP/CBC) picks where to place sports centres, health centres or bike stations to maximize population coverage under a real budget.
The interesting part isn't the model — it's everything around it. A model doesn't end at evaluation: it has to ship as a service, publish itself automatically, and be observable once it's running. So the repository implements the full CRISP-DM cycle with the emphasis on the phases academic projects usually skip. FastAPI on Hugging Face, Next.js on Vercel, CI/CD through GitHub Actions, and a monitoring module that computes MAE per hour, raises alerts past a configurable threshold, and flags zones with systematic error — so degradation surfaces without anyone inspecting it by hand.
Validation is deliberately strict (temporal hold-out on days 25–31, never a random split), and the app pulls live data from Valenbisi, EMT buses, City Council traffic status and AEMET/Open-Meteo weather.
🔗 Live demo · API docs
Python · FastAPI · CatBoost · PuLP/CBC · Next.js · TypeScript · Docker · GitHub Actions · Vercel · Hugging Face
Extreme congestion physically blocks emergency vehicles, and the cost is measurable: the golden minutes are lost, a trapped unit is a failed emergency, and fleet availability drops as the bottleneck propagates. The founding idea is one sentence — predict traffic so that emergency services don't have to deal with traffic jams.
The gap AION fills is that commercial navigation apps are reactive: they see traffic only while the app is open, and they're blind to a concert that ends at 23:00 or roadworks announced yesterday. AION instead reads the city's recorded sensor history and models the road itself — type, speed limit, direction, width, lane count — on the premise that raw intensity isn't comparable across streets. It predicts per zone × hour at a 30–120 minute horizon, targeting deviation from normality rather than absolute volume.
A robust baseline paired with PCA embeddings, CatBoost and a sigmoid shrink step — which keeps it stable in low-signal hours without flattening the peaks — reaches R² 0.971 · MAE 44.5 veh/h · RMSE 93.9 · sMAPE 17.4%. Feeding it is an LLM pipeline (LLaMA 3 via Ollama) that filters scraped local news, city agendas and social posts for traffic relevance, then turns the survivors into structured events with severity, cause, entities, time window and geocoding — the part that lets the model know Fallas is starting.
Capstone of the University Extension Diploma in AI (467h) — Samsung Innovation Campus · EOI.
Python · CatBoost · LLaMA 3 / Ollama · MongoDB · Streamlit · scraping & ETL
Outfit AI Recommender — multimodal fashion discovery
Winner of the Retail & Fashion challenge at Accenture Spain's GenAI Maverick 2025. Finds catalogue items from a sentence, a garment photo, a full outfit (segmented and matched piece by piece) or just a situation — "an outdoor wedding in June" — which isn't a garment description at all and needs its own path through the system.
FastAPI · FAISS · BLIP · SegFormer
GENAQ — market selection for atmospheric water generators
Ranks Spain's 35,891 census sections to find the 50 best places to launch a machine that makes drinking water out of air. The catch: no label exists — nobody has sold this to a Spanish household — so the target has to be argued from first principles rather than trained. Bankinter Akademia 2025 case study for GENAQ.
Python · INE census & climate data · composite scoring
ExamBox — self-hosted Docker exam platform
Runs practical programming exams on a LAN by making the environment part of the exam: a pinned Docker image, so no install fails mid-exam and no submission has to be reconstructed to grade it. FastAPI manager with live dashboard and progressive question unlocking; students get browser-based JupyterLab and zero setup. Graded 10/10.
Python · FastAPI · Docker Compose · JupyterLab
All 508 players of the 2022–23 season through PCA + k-means to find who can replace whom, and PLS-DA that catches 100% of All-Stars at 95% specificity. The best result is a failure: its "false" positives are players the stats say deserved selection but the fan vote didn't give it to.
R · R Markdown · PCA · k-means · PLS-DA
Six visualizations over a 189-country × 366-day panel, and an uncomfortable 2020 answer: wealth correlates with more deaths per capita (ρ +0.49), not fewer — largely a detection effect, since richer countries tested more. Built to generate hypotheses, not settle them.
Python · Shiny for Python · Plotly
NOBIL Data — real-time EV charging pipeline
Websocket ingestion of NOBIL charging events into hourly-rotated JSONL with periodic snapshots, derived real-time metrics (status transitions, dwell time, event lag, duplicate detection) and automated versioned sync to GitHub and Google Cloud Storage.
Python · websockets · Git automation · GCS
Streamlit app answering "what do I train, what do I eat, and what should I buy?" from one form — Harris-Benedict calorie targets, 112 scraped routines, weekly menus with macros via Spoonacular, and supplements fuzzy-matched against a 5,023-product Decathlon catalogue.
Python · Streamlit · BeautifulSoup · REST APIs
I am particularly interested in systems that connect data, models, infrastructure, and deployment-oriented thinking. Beyond model development, I care about maintainability, reproducibility, automation, cloud workflows, and operational value, which is why I am increasingly drawn to MLOps and DevOps practices.
I also actively use modern AI tooling to accelerate experimentation, coding workflows, and problem solving, especially across the broader LLM ecosystem.
- Go for backend and systems-oriented development
- Cloud workflows with Google Cloud and AWS
- MLOps / DevOps practices for reproducible and deployable ML systems
- AI-assisted engineering workflows with tools such as ChatGPT, Claude, Gemini, Claude Code, and GitHub Copilot
Turning complexity into useful solutions.




