STATlee is AI data analysis for social scientists. Upload a dataset, describe the analysis in plain English, and STATlee writes, moderates, sandboxes, runs, and explains the statistics. No Python, no R, no syntax errors.
Architecture · Pricing · Credits
Upload data, ask in plain English, get runnable code, a chart, and a written report. Try it with the included sample survey dataset.
Statistical software is hostile to non-coders. SPSS menus, Stata syntax, and R tibbles all stand between a researcher and a simple question like "does income predict turnout?" STATlee removes that wall: you ask in plain English and get a runnable analysis, charts, and a plain-English write-up you can defend.
Try it: the marketing landing page lives at
/welcome; the app itself is at/.
- Describe it, don't code it. Natural-language requests become real Python/R that runs and returns results, not just a code snippet.
- Statistically valid by design. An intelligent codebook classifies each variable (nominal / ordinal / continuous) so the model won't run a linear model on a categorical outcome. Codebooks can be inferred from a PDF data dictionary or even the original survey questionnaire.
- Secure sandbox. Generated code runs in a throwaway, network-less
working directory with a secret-free environment;
SANDBOX_MODE=dockeradds a non-root, read-only, resource-capped container per run. A run-guard re-moderates any hand-edited script before it executes. - Answers you can explain. Dense terminal output and p-values become plain-English Markdown (effect sizes, significance, caveats), and a debugging assistant kicks in when a run fails.
- Bring any format. CSV, TSV, Excel (
.xlsx/.xls), Stata (.dta), and SPSS (.sav); native value labels seed the codebook for free. - Conversational wrangling. Clean your data by chatting ("delete the
notes column", "filter for age > 30") with a back-and-forth transcript, full
version history (undo / redo), and a one-click revert to the original
upload. Runs on the cheapest model tier (
WRANGLE_ROLE=lite) to keep costs down. Plus an AI report builder grounded strictly in your real outputs and one-click project export (data + script + plots + report). - Pro mode toggle generates code with a larger model (
gemini-3.1-pro-previewby default) for hard or unusual analyses, instead of the standard code model.
Upload data ──▶ Intelligent codebook ──▶ Describe analysis (plain English)
│ │
▼ ▼
Multi-format moderation ▶ feature-select ▶ draft ▶ validate
normalize to CSV │
▼
Secure sandbox (subprocess / Docker)
│
▼
Plain-English interpretation + charts ──▶ Report / export
Every model call goes through one role-based LLM service (pro / flash /
lite / draft), so swapping a model, switching LLM provider, or pointing Pro
mode at a stronger model is a config change, not a code change. Per-request
token usage is surfaced live in the UI.
cp .env.example .env # add your GEMINI_API_KEY
docker-compose up --build # then open http://localhost:5000Or run locally without Docker (dev only, generated code runs on your host):
python -m venv venv && source venv/bin/activate # Windows: venv\Scripts\activate
pip install -r requirements.txt
APP_ENV=development python wsgi.py # http://localhost:5000STATlee has a pluggable LLM provider: pick one with LLM_PROVIDER and
supply that vendor's key. Switching providers is a one-line config change; no
route, prompt, or UI code changes (every call routes through one role-based
service).
LLM_PROVIDER |
API key | Notes |
|---|---|---|
gemini (default) |
GEMINI_API_KEY (AI Studio) |
Powers the multimodal paths (PDF codebook extraction, plot interpretation). |
anthropic |
ANTHROPIC_API_KEY (Console) |
Claude models; also accepts a local OAuth/subscription login for keyless dev. |
openai |
OPENAI_API_KEY (Platform) |
GPT models. |
The selected provider's key is required in production. Each role's model id
defaults sensibly per provider and can be pinned independently with MODEL_PRO
/ MODEL_PRO_MAX / MODEL_FLASH / MODEL_FLASH_LITE (see
.env.example for the per-provider defaults).
See docs/README.md for the full setup, Docker sandbox build,
and the documented .env.example.
pip install -r requirements-dev.txt
ruff check . # lint
pytest -q # 334 tests, fully offline (deterministic fake LLM, no API key)The test suite injects a fake LLM service, so the entire HTTP surface (uploads,
codebook, wrangling, run-guard, converse, export, auth, CSRF, rate-limit keying,
billing seam, Pro-mode routing) is exercised without network or API keys. CI runs
ruff + byte-compile + pytest on every push (.github/workflows/ci.yml).
The application lives in the statlee/ package (entry point: wsgi.py →
statlee.app:app).
| Module | Responsibility |
|---|---|
statlee/config.py |
Validated, env-driven configuration (one source of truth). |
statlee/app.py |
App factory + middleware (sessions, CSRF, rate limits, ProxyFix, request-id logging). |
statlee/storage.py |
Per-identity file isolation + dataset version control. |
statlee/sandbox.py |
Isolated code execution (subprocess or Docker). |
statlee/llm.py |
Role-based LLM service (pluggable Gemini / Claude / OpenAI backend): usage tracking, Pro-mode routing, response cache. |
statlee/billing.py |
Monetization seam: check_and_debit chokepoint (off by default; atomic debit + refund + monthly top-up when enabled). |
statlee/prompts.py |
Every prompt builder in one reviewable place. |
statlee/datatools.py |
Multi-format ingestion + metadata profiling. |
statlee/models.py |
SQLAlchemy models (users, runs, issue reports). |
statlee/routes/ |
Blueprints: auth, datasets, analyze, converse, misc. |
A deeper walkthrough (request lifecycle, security boundaries, data flow) is in docs/ARCHITECTURE.md.
STATlee runs locally for free and is not currently deployed to a public website, a deliberate, zero-cost choice. Self-hosted, its only real cost is the LLM provider's API (Gemini by default), and it ships with the controls to cap that spend at a number you choose: a global monthly ceiling on premium calls, per-identity rate limits, cheapest-tier defaults, and a startup warning if billing is on without a cap. The (illustrative) monetization model and how the billing seam backs it are in Pricing.
The execution sandbox is the real security boundary: secret-free env, throwaway dir, optional network-less container, plus a run-guard that re-moderates edited scripts. Rate limits are keyed per-identity (client IP / account, not a resettable cookie) to protect against bill abuse. Responses carry a strict Content-Security-Policy and the usual hardening headers, CSRF is double-submit with a constant-time check, and the schema is versioned with Alembic so an existing database upgrades cleanly instead of breaking on a new column.
Found a vulnerability? See SECURITY.md for how to report it privately. Release notes are in CHANGELOG.md.
STATlee's AI analysis is powered by your configured LLM provider, Google Gemini by default (with optional Anthropic Claude or OpenAI backends), and the in-app footer attributes the active provider. Default Gemini use is subject to the Gemini API Additional Terms and the Prohibited Use Policy, which STATlee's moderation gate enforces.
Third-party libraries and their licenses are listed in CREDITS.md.
STATlee is a research aid: always review generated code and results before relying on them.
STATlee is source-available under the Elastic License 2.0: you're free to use, modify, and self-host it, but you may not offer it to third parties as a hosted or managed service. This keeps the source open to read and learn from while reserving the right to run it commercially. See CREDITS.md for third-party dependency licenses.


