Skip to content

Add a nine-notebook tutorial series, from first model to contributing - #1281

Open
solarsys wants to merge 6 commits into
sunlabuiuc:masterfrom
solarsys:docs/tutorial-series
Open

solarsys wants to merge 6 commits into
sunlabuiuc:masterfrom
solarsys:docs/tutorial-series

Conversation

@solarsys

@solarsys solarsys commented Oct 9, 2026 •

Copy link
Copy Markdown
Collaborator

Why

The Colab tutorials in the shared Drive folder and on the docs page have drifted from the code. I ran the five in the Drive folder against PyHealth 2.0.2, the version Colab installs today:

Notebook Result on 2.0.2
Tutorial 1 (datasets) Runs. Its class walkthrough shows polars, but preprocess_<table> hooks now receive narwhals frames.
Tutorial 2 (tasks) Runs. Needs an 816 MB download for the X-ray part, and trains Bio_ClinicalBERT on CPU.
MIMIC3MortalityPrediction Pins pyhealth==2.0a4. A recursive PhysioNet download takes about 10 minutes and yields 26 samples, so validation ROC-AUC is 1.0 after one epoch. It splits by sample, not by patient.
MedicalCoding Pins 2.0a4. About 5 minutes per CPU epoch. Never says the synthetic notes carry no signal.
MedicalTranscriptions Pins 2.0a4. 24 minutes for one CPU epoch, ending at macro-F1 0.009 (it predicts the majority class), with no baseline to reveal it.

None of them covers patient-level splitting without leakage, imbalance, calibration, interpretation or medical codes.

What this adds

examples/tutorials/series/ contains nine notebooks, in order of complexity. Each runs on a free Colab CPU runtime with public or synthetic data, so no credentials are needed.

# Notebook CPU time* Key point it teaches
00 Quickstart ~2 min Dataset → task → split → model → evaluate; PR-AUC against its no-skill baseline
01 Datasets ~1 min Events and filters, the YAML config, caching (PYHEALTH_REQUIRE_CACHE_DIR), loading your own CSVs with BaseDataset
02 Tasks & processors ~5 min What each processor produces, custom tasks, pre_filter, leak-free features
03 Training & evaluation ~2.5 min set_task(split=PatientSplit(...)) (5% of test codes are unseen in training), set_pos_weight inflates risk (0.05 → 0.45), calibration metrics, patient-level bootstrap_ci / paired_bootstrap_diff, checkpoints
04 Choosing a model ~4 min LR and MLP (+ RNN, Transformer and RETAIN on GPU) vs XGBoostModel on one split. On Colab CPU: XGBoost PR-AUC 0.478, MLP 0.454, LR 0.384
05 Interpreting predictions ~2 min Exact TreeSHAP (furosemide ranks first; additivity check), Integrated Gradients, and a deletion test (IG beats random roughly 3×)
06 Medical codes ~5 min InnerMap / CrossMap, ICD-9↔ICD-10 not being a round trip, CCS/ATC groupers, code_mapping (2,528 → 257 diagnosis groups)
07 Clinical text ~3.5 min TransformersModel sized to the hardware (Bio_ClinicalBERT for 3 epochs on GPU, small BERT for 1 epoch on CPU), a TF-IDF baseline, ICD coding as a multilabel task
08 Contributing <1 min Dataset class hooks (preprocess_<table> with narwhals, default_task), the metadata-table pattern for file collections, inline tests, the PR checklist

* Measured on CPU (macOS, Python 3.13), excluding the install cell.

Phenotyping heart failure from an admission's medications (04, 05) is the task with real signal in the synthetic data; readmission and mortality are close to random there, and the notebooks say so where they use them.

  • Single source of truth: the notebooks are generated from _source/tutorials/tNN_*.py by _source/build.py, so wording and code stay consistent and diffs stay readable. Cell ids are stable, and --run executes a notebook and stops at the first error. README.md explains how to edit them.
  • Docs: docs/tutorials.rst lists the series first, with Colab links that open the notebooks straight from GitHub (colab.research.google.com/github/...). The older list stays below under "Earlier tutorials", so you can retire it when you like.
  • Drive copies: the same nine notebooks are in a new subfolder of the shared tutorials folder, "PyHealth Tutorials (2.1 draft)". The originals there are untouched.

Notes

  • Install line: until 2.1 is on PyPI, the notebooks install from GitHub. After the release, the install cell becomes pip install "pyhealth>=2.1"; that's one line in _source/nbkit.py, then a rebuild.
  • Tested on Colab: all nine notebooks now run end to end on a free Colab CPU runtime (Python 3.13), as well as locally. Getting there needed these changes:
    • Self-restarting install cell. PyHealth's pins (numpy~=2.2, pandas~=2.3.1, pydantic~=2.11.7) downgrade packages Colab has already imported, so the first import pyhealth failed ('numpy.ufunc' object has no attribute '__module__') and pip printed resolver errors. The cell now installs quietly, restarts the runtime once, and skips the install on the second Run all. Colab shows a "session crashed" notice at the restart; the Setup text says this is expected.
    • Logging printed every line twice. The pyhealth logger has its own stdout handler and also propagates to the root logger, which Colab prints too. The install cell sets propagate = False.
    • Runtime on CPU. Tutorial 04 trains the sequence models only on a GPU (RNN ran at about 3 s/iteration on Colab CPU). Tutorial 07 trains the small BERT for 1 epoch on CPU and recommends a GPU up front.
    • Tutorial 08 now prints the files it creates instead of a character count.
  • Issues found while writing these:
    • pyhealth.metrics.interpretability: the "zero" ablation multiplies code-id tensors by a float mask, so comprehensiveness and sufficiency crash for every model with code-sequence inputs; only StageNet's tuple inputs take the integer-safe path. Tutorial 05 uses a hand-written deletion test instead. This needs a small fix PR.
    • Outside Colab, a cold-cache dataset build in a notebook needs ipywidgets, which Build datasets in notebooks without ipywidgets; fix datasets overview example #1260 fixes.
    • split_by_patient does not sort patient ids before shuffling, so the same seed gives different splits in different sessions. PatientSplit does sort.
    • The numpy/pandas/pydantic pins above conflict with Colab's preinstalled versions. Loosening them would remove the restart.
    • The duplicate-logging issue above is better fixed in the library than in each notebook.
    • In synthetic MIMIC-III every birth date sits just before the first admission, so ReadmissionPredictionMIMIC3's default exclude_minors=True returns no samples. Tutorial 02 explains this and turns the option off.

🤖 Generated with Claude Code

The Colab tutorials linked from the docs had drifted from the code: three
pin pyhealth==2.0a4, one trains on 26 samples from the PhysioNet demo
(validation ROC-AUC 1.0) and splits by sample, the text ones train
Bio_ClinicalBERT on CPU for ~25 minutes per epoch without a baseline, and
none covers patient-level splitting, calibration, interpretation or
medical codes.

examples/tutorials/series/ adds nine notebooks with a progression of
complexity, each verified end to end on CPU against master with public or
synthetic data:
00 quickstart, 01 datasets, 02 tasks and processors, 03 training and
evaluation (PatientSplit, pos_weight, calibration, patient bootstrap CIs),
04 choosing a model (incl. XGBoost), 05 interpreting predictions (TreeSHAP,
Integrated Gradients, deletion test), 06 medical codes, 07 clinical text,
08 contributing a dataset or task.

The notebooks are generated from _source/tutorials/*.py by
_source/build.py (stable cell ids; --run executes them). docs/tutorials.rst
lists the series first, with Colab links that open from GitHub.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@solarsys
solarsys requested a review from jhnwu3 October 9, 2026 11:17
solarsys and others added 5 commits October 9, 2026 10:24
On Colab, installing PyHealth downgrades numpy (and pandas, pydantic)
under the copies Colab already loaded, so the first import failed with
"'numpy.ufunc' object has no attribute '__module__'", and pip printed a
red dependency-resolver report. The install cell now installs quietly
(showing pip's log only on failure), restarts the runtime once, and skips
the install on the next Run all. Verified on Colab.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
PyHealth's logger has its own stdout handler and also propagates to the
root logger, which Colab configures, so every message printed twice.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
On Colab's free 2-CPU runtime the RNN on the heart-failure task needs
~7 minutes per epoch (long medication lists); sequence models now run
only on a GPU runtime.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
07 trains one epoch on a CPU runtime (~10 minutes per epoch on Colab's
free 2-CPU runtime) and recommends a GPU runtime up front; 04 no longer
claims trees train much faster than the pooled neural models, which is
not true on 2 CPUs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant