AI Product / Program Manager | Independent AI Researcher | Human-Centered AI & Evaluation
PhD in Art Therapy with deep expertise in human behavior, emotional processing, and clinical research pipelines. Former Senior Program Manager / Knowledge Manager at Amazon (Middle Mile Transportation), scaling internal knowledge systems for 900+ people (250+ engineers, ~50 scientists).
Currently building modular AI evaluation frameworks, data pipelines from unstructured sources (books/memoirs/images), and metrics-driven tools for trustworthy, empathetic AI.
- Location: Redmond, WA (Seattle area) | Open to hybrid/remote AI PM, TPM, Applied Scientist roles (L4/L5)
- Email: limorgu@gmail.com
- LinkedIn: linkedin.com/in/limorkissos
AI-Powered Book Dataset Builder
Automates turning raw book page images into clean, structured, research-ready JSON datasets.
- Three-stage pipeline: extraction → audit for gaps → precise gap-filling
- Resume-aware, cost-efficient, and audit-driven
- Ideal for large-scale qualitative research from physical books
Modular Multi-Stage AI Pipeline for Research Datasets
Full end-to-end framework for building structured datasets from physical books.
- Stages include workspace setup, librarian intake (OCR), worker extraction, judge audit, analytics, reports, ground-truth export, and benchmarking
- Highly configurable via JSON (domains, taxonomy, labels)
- Supports deterministic/local and future LLM connectors
Raw Data → Meaningful Insights Pipeline
Stable, Codex-compatible version of the book processing workflow.
- Turns raw inputs into categorized insights using configurable taxonomies
- Includes full stage orchestration and reporting
MASK Honesty Benchmark Evaluation (Fork)
Implementation for evaluating AI honesty (disentangling it from accuracy) using the MASK benchmark.
- Tests model consistency under pressure to lie
- nature-ai-pipeline — AI pipeline work focused on natural domains/data
- comparing_architecture_classification- — Architecture comparison & classification experiments
-
Amazon (2021–2025): Senior Knowledge Manager / Program Manager
Built and scaled data-driven knowledge & workflow systems for large ops teams. Designed metrics/experimentation frameworks, led AWS enablement workshops (900+ participants), science newsletter. -
Independent Research (2023–Present):
LLM evaluation (empathy, alignment, grounding, sycophancy via ELEPHANT-style studies), synthetic data generation, narrative/therapy AI tools, and memoir dysregulation analysis. -
Academic Background: PhD Art Therapy (University of Haifa), published research on AI detection of childhood sexual abuse in drawings (~72% accuracy).
Product & Program Management — Strategy, experimentation (A/B), metrics & dashboards, cross-functional leadership, user research
AI/ML — Evaluation frameworks, RAG/pipelines, synthetic data, LLM prompting & evaluation
Technical — Python, Pandas, SQL, JSON data pipelines, Git, API integration
Last updated: June 2026
