Skip to content

Data-frame API, theory, standard errors; release 0.5.0 - #13

Merged
soodoku merged 7 commits into
mainfrom
theory-and-priorities
Sep 30, 2026
Merged

soodoku merged 7 commits into
mainfrom
theory-and-priorities

Conversation

@soodoku

@soodoku soodoku commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Release 0.5.0. Breaking changes are listed in CHANGELOG.md.

  • Public API: calibrate(df, targets, ...), calibrate_replicates, CalibrationReport. The report names the worst-missed targets and says whether other targets or the weight bounds are to blame.
  • Theory page: the largest relative miss equals the worst-case calibration bias over a stated class of outcomes. Each claim is checked in tests/test_theory.py.
  • Standard errors page: recalibrate replicates. The simcheck study shows Rao-Wu understates SEs about threefold when bounds bind, while a Poisson bootstrap is calibrated.
  • Verification: closed-form cases, Hypothesis properties, and independent oracles for both stages. Each is shown to fail on a broken solver.
  • Requires pandas >= 2.1.1.

Local checks: make ci (165 passed, 98.8% coverage), make docs, preen, twine; wheel install plus README example on Python 3.14; tests on numpy 1.26.0 / scipy 1.12.0 / pandas 2.1.1.

🤖 Generated with Claude Code

- scale accepts one positive number per margin; margin j's miss is
  divided by s_j, so users state which margins matter.
- docs/theory.rst: epsilon equals the worst-case calibration bias over
  outcomes with sum_j s_j |beta_j| <= 1 (Holder), leximin as its
  lexicographic refinement, properties and limits.
- tests/test_theory.py checks each claim by an independent route.
- README: "fair" means sharing misses across margins, not ML fairness.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-30T23:27:36.380046Z 0ed17b8 New commits
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

Lead with the situation fairlex is for (targets that cannot all be hit),
show the value on a conflicting-targets example whose printed numbers
come from running it, and say when not to use it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 6bcd400e10

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread docs/theory.rst
Comment on lines +34 to +36
The first term is the part calibration controls. The second depends on how
well the margins explain :math:`y`, which no choice among these weights can
change.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Describe the remainder as uncontrolled rather than invariant

The remainder term explicitly contains the selected weights, so it can change whenever two feasible calibrated solutions assign different weights to units with nonzero e_i. This is particularly relevant because leximin_weight_fair deliberately selects among residual-optimal weights. The margin constraints do not bound this term, but saying no choice of weights can change it overstates the guarantee.

Useful? React with 👍 / 👎.

Comment thread docs/theory.rst
Comment on lines +77 to +79
Adding a total row alongside categories that sum to it is not double
counting. It adds outcomes that load on the total to the protected set. Leave
it out if you do not want that.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Warn that a redundant total can change the objective

When the total row is the sum of the category rows, it adds no new outcome direction but does add another residual to the leximin objective and can therefore change the weights. For example, under absolute scaling, if two category residuals are x and a fixed 1, omitting the total makes the refinement choose x=0, while adding its residual x+1 makes it choose x=-1/2. Calling this “not double counting” can lead users to add a redundant row and unintentionally reweight the same discrepancy.

Useful? React with 👍 / 👎.

benchmarks/regimes.py compares fairlex with svy raking-plus-trimming and
ridge calibration when calibration targets disagree and weights are
capped (30 samples x 6 regimes, with standard errors). The docs page
reads its summary CSV directly. Adds a 'bench' dependency group.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 80924b0acc

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/fairlex/calibration.py Outdated
Comment on lines +174 to +175
s = np.asarray(scale, dtype=float)
if s.shape != (m,) or not np.all(np.isfinite(s)) or np.any(s <= 0):

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Normalize custom scales before applying fixed LP tolerances

When a custom scale is much larger than its target, passing it through unchanged makes the fixed _TOL floor correspond to a large raw-residual allowance. For example, with one reachable margin whose base total is 1000, target is 1100, and scale=[1e9], the weight-stage holding band is about 300 raw units, so with slack=0 the solver can select the unchanged base weights with a miss of 100 instead of preserving the exact stage-one solution. Consequently, multiplying every custom scale by a common factor can change the returned weights even though the leximin priorities are unchanged; normalize the scale used internally or make the holding tolerance invariant to this common rescaling.

Useful? React with 👍 / 👎.

Comment thread README.md
Comment on lines +64 to +65
If your targets are consistent and reachable within your weight caps, every calibration
method hits them, and the choice among methods matters less. Standard raking in

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Restrict the exact-hit claim to exact calibration methods

Consistent, reachable targets do not imply that every calibration method hits them: a finite-penalty method can trade residual error for smaller weight changes. The benchmark added in this commit demonstrates this at zero target noise and a 4x cap, where ridge with lambda 10 has a 1.63% worst miss and ridge with lambda 1000 still has a nonzero 0.031% miss while fairlex reaches the targets. This guidance should be limited to exact calibration methods rather than telling users that method choice is immaterial in this regime.

Useful? React with 👍 / 👎.

- Public API is calibrate, calibrate_replicates and CalibrationReport.
  The matrix functions, string scales and evaluate_solution are gone;
  the engine is fairlex.calibration.leximin_weights.
- CalibrationReport names the worst-missed targets and whether other
  targets or the weight bounds are to blame.
- Standard errors: calibrate_replicates, with a simcheck-validated study
  (benchmarks/se_study.py) and a docs page. Rao-Wu understates SEs
  about threefold when bounds bind; a Poisson bootstrap does not.
- Verification: closed-form cases, Hypothesis properties, and
  independent oracles for both leximin stages, each shown to fail on a
  deliberately broken solver.
- Drop the regime study. Require pandas >= 2.1.1; tests pass on
  numpy 1.26.0, scipy 1.12.0, pandas 2.1.1.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@soodoku soodoku changed the title Per-margin scale priorities and a theory page Data-frame API, theory, standard errors; release 0.5.0 Sep 30, 2026

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 963d44d166

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/fairlex/frame.py
at_bound = np.all(
np.abs(ratio[members] - helpful) <= _AT_BOUND_RTOL * helpful
)
reason = "weight bounds" if at_bound else "conflicting targets"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Distinguish slack-selected misses from target conflicts

When slack > 0, the weight-refinement stage may deliberately leave an otherwise feasible margin away from its target to reduce weight changes. For example, a single respondent with base weight 1, target 2, and slack=0.5 can remain at weight 1; because the respondent is not at a bound, this line reports "conflicting targets" even though no other target exists. Compare against the pre-slack leximin levels or report slack as a separate reason instead of diagnosing every non-bound miss as a conflict.

Useful? React with 👍 / 👎.

Comment thread src/fairlex/frame.py Outdated
Comment on lines +194 to +197
if size is None:
size = float(
np.mean([b[[v == var for v in variables]].sum() for var in targets])
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Require an explicit size for inconsistent target totals

When base_weight and total are both omitted and different variables imply different population sizes, this silently uses their mean despite the API documentation saying total is required. That inferred size determines every base weight and therefore the feasible interval under ratio bounds, so targets summing to 100 and 200 are calibrated relative to an arbitrary total of 150 rather than producing the promised validation error. Detect unequal target sums and require total in this branch.

Useful? React with 👍 / 👎.

soodoku and others added 3 commits September 30, 2026 16:18
From independent review of the 0.5.0 candidate:
- a miss allowed by slack was reported as 'conflicting targets';
- a missing total with disagreeing variable sizes silently used the mean;
- shares not summing to 1 were silently renormalised;
- negative targets 'succeeded' with a sign-flipped miss_pct.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Shares within 0.01 of summing to 1 are rescaled to sum exactly to 1, so
0.333 x 3 no longer produces false conflicts and 0.998 is not rejected.
Found in independent review.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A Hypothesis case showed the oracle, holding margins slightly looser than
the engine, reaching changes 1.2e-4 smaller on an ill-conditioned problem.
Both now solve the same problem, and the 1e-4 tolerance holds.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@soodoku
soodoku merged commit 0ac66a0 into main Sep 30, 2026
13 checks passed
@soodoku
soodoku deleted the theory-and-priorities branch September 30, 2026 23:26

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 0ed17b862c

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/fairlex/frame.py
msg = "shares=True needs total= to turn proportions into totals"
raise ValueError(msg)
sums = _variable_sums(b, variables)
off = {v: s for v, s in sums.items() if abs(s - 1) > _SHARE_TOL}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Scale the share tolerance by the number of levels

Use a level-count-dependent rounding tolerance here rather than the fixed 0.01. With a two-level target such as 0.495 and 0.496, the sum is 0.991—far outside what rounding each share to three decimals can explain—but this check accepts it and silently rescales the shares to sum to one. This can conceal malformed inputs and change targets by nearly a percentage point despite the API promising to reject sums beyond rounding.

Useful? React with 👍 / 👎.

Comment thread src/fairlex/frame.py
Comment on lines +128 to +133
if not float(value) >= 0:
msg = (
"targets must be non-negative numbers, "
f"got {var}={level!r}: {value!r}"
)
raise ValueError(msg)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Handle zero-valued targets before deriving relative scales

Either reject zero here as requiring strictly positive targets or define a usable scale for it. This check explicitly accepts zero, and the public documentation likewise permits non-negative targets, but calibrate later passes abs(b) / factor to leximin_weights, whose positive-scale validation rejects every zero target. Thus a represented category with a zero total or share always fails despite passing this front-end validation, even when bounds=(0, ...) would make the target reachable.

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant