Data-frame API, theory, standard errors; release 0.5.0 - #13
Conversation
- scale accepts one positive number per margin; margin j's miss is divided by s_j, so users state which margins matter. - docs/theory.rst: epsilon equals the worst-case calibration bias over outcomes with sum_j s_j |beta_j| <= 1 (Holder), leximin as its lexicographic refinement, properties and limits. - tests/test_theory.py checks each claim by an independent route. - README: "fair" means sharing misses across margins, not ML fairness. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
Lead with the situation fairlex is for (targets that cannot all be hit), show the value on a conflicting-targets example whose printed numbers come from running it, and say when not to use it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 6bcd400e10
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| The first term is the part calibration controls. The second depends on how | ||
| well the margins explain :math:`y`, which no choice among these weights can | ||
| change. |
There was a problem hiding this comment.
Describe the remainder as uncontrolled rather than invariant
The remainder term explicitly contains the selected weights, so it can change whenever two feasible calibrated solutions assign different weights to units with nonzero e_i. This is particularly relevant because leximin_weight_fair deliberately selects among residual-optimal weights. The margin constraints do not bound this term, but saying no choice of weights can change it overstates the guarantee.
Useful? React with 👍 / 👎.
| Adding a total row alongside categories that sum to it is not double | ||
| counting. It adds outcomes that load on the total to the protected set. Leave | ||
| it out if you do not want that. |
There was a problem hiding this comment.
Warn that a redundant total can change the objective
When the total row is the sum of the category rows, it adds no new outcome direction but does add another residual to the leximin objective and can therefore change the weights. For example, under absolute scaling, if two category residuals are x and a fixed 1, omitting the total makes the refinement choose x=0, while adding its residual x+1 makes it choose x=-1/2. Calling this “not double counting” can lead users to add a redundant row and unintentionally reweight the same discrepancy.
Useful? React with 👍 / 👎.
benchmarks/regimes.py compares fairlex with svy raking-plus-trimming and ridge calibration when calibration targets disagree and weights are capped (30 samples x 6 regimes, with standard errors). The docs page reads its summary CSV directly. Adds a 'bench' dependency group. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 80924b0acc
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| s = np.asarray(scale, dtype=float) | ||
| if s.shape != (m,) or not np.all(np.isfinite(s)) or np.any(s <= 0): |
There was a problem hiding this comment.
Normalize custom scales before applying fixed LP tolerances
When a custom scale is much larger than its target, passing it through unchanged makes the fixed _TOL floor correspond to a large raw-residual allowance. For example, with one reachable margin whose base total is 1000, target is 1100, and scale=[1e9], the weight-stage holding band is about 300 raw units, so with slack=0 the solver can select the unchanged base weights with a miss of 100 instead of preserving the exact stage-one solution. Consequently, multiplying every custom scale by a common factor can change the returned weights even though the leximin priorities are unchanged; normalize the scale used internally or make the holding tolerance invariant to this common rescaling.
Useful? React with 👍 / 👎.
| If your targets are consistent and reachable within your weight caps, every calibration | ||
| method hits them, and the choice among methods matters less. Standard raking in |
There was a problem hiding this comment.
Restrict the exact-hit claim to exact calibration methods
Consistent, reachable targets do not imply that every calibration method hits them: a finite-penalty method can trade residual error for smaller weight changes. The benchmark added in this commit demonstrates this at zero target noise and a 4x cap, where ridge with lambda 10 has a 1.63% worst miss and ridge with lambda 1000 still has a nonzero 0.031% miss while fairlex reaches the targets. This guidance should be limited to exact calibration methods rather than telling users that method choice is immaterial in this regime.
Useful? React with 👍 / 👎.
- Public API is calibrate, calibrate_replicates and CalibrationReport. The matrix functions, string scales and evaluate_solution are gone; the engine is fairlex.calibration.leximin_weights. - CalibrationReport names the worst-missed targets and whether other targets or the weight bounds are to blame. - Standard errors: calibrate_replicates, with a simcheck-validated study (benchmarks/se_study.py) and a docs page. Rao-Wu understates SEs about threefold when bounds bind; a Poisson bootstrap does not. - Verification: closed-form cases, Hypothesis properties, and independent oracles for both leximin stages, each shown to fail on a deliberately broken solver. - Drop the regime study. Require pandas >= 2.1.1; tests pass on numpy 1.26.0, scipy 1.12.0, pandas 2.1.1. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 963d44d166
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| at_bound = np.all( | ||
| np.abs(ratio[members] - helpful) <= _AT_BOUND_RTOL * helpful | ||
| ) | ||
| reason = "weight bounds" if at_bound else "conflicting targets" |
There was a problem hiding this comment.
Distinguish slack-selected misses from target conflicts
When slack > 0, the weight-refinement stage may deliberately leave an otherwise feasible margin away from its target to reduce weight changes. For example, a single respondent with base weight 1, target 2, and slack=0.5 can remain at weight 1; because the respondent is not at a bound, this line reports "conflicting targets" even though no other target exists. Compare against the pre-slack leximin levels or report slack as a separate reason instead of diagnosing every non-bound miss as a conflict.
Useful? React with 👍 / 👎.
| if size is None: | ||
| size = float( | ||
| np.mean([b[[v == var for v in variables]].sum() for var in targets]) | ||
| ) |
There was a problem hiding this comment.
Require an explicit size for inconsistent target totals
When base_weight and total are both omitted and different variables imply different population sizes, this silently uses their mean despite the API documentation saying total is required. That inferred size determines every base weight and therefore the feasible interval under ratio bounds, so targets summing to 100 and 200 are calibrated relative to an arbitrary total of 150 rather than producing the promised validation error. Detect unequal target sums and require total in this branch.
Useful? React with 👍 / 👎.
From independent review of the 0.5.0 candidate: - a miss allowed by slack was reported as 'conflicting targets'; - a missing total with disagreeing variable sizes silently used the mean; - shares not summing to 1 were silently renormalised; - negative targets 'succeeded' with a sign-flipped miss_pct. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Shares within 0.01 of summing to 1 are rescaled to sum exactly to 1, so 0.333 x 3 no longer produces false conflicts and 0.998 is not rejected. Found in independent review. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A Hypothesis case showed the oracle, holding margins slightly looser than the engine, reaching changes 1.2e-4 smaller on an ill-conditioned problem. Both now solve the same problem, and the 1e-4 tolerance holds. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 0ed17b862c
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| msg = "shares=True needs total= to turn proportions into totals" | ||
| raise ValueError(msg) | ||
| sums = _variable_sums(b, variables) | ||
| off = {v: s for v, s in sums.items() if abs(s - 1) > _SHARE_TOL} |
There was a problem hiding this comment.
Scale the share tolerance by the number of levels
Use a level-count-dependent rounding tolerance here rather than the fixed 0.01. With a two-level target such as 0.495 and 0.496, the sum is 0.991—far outside what rounding each share to three decimals can explain—but this check accepts it and silently rescales the shares to sum to one. This can conceal malformed inputs and change targets by nearly a percentage point despite the API promising to reject sums beyond rounding.
Useful? React with 👍 / 👎.
| if not float(value) >= 0: | ||
| msg = ( | ||
| "targets must be non-negative numbers, " | ||
| f"got {var}={level!r}: {value!r}" | ||
| ) | ||
| raise ValueError(msg) |
There was a problem hiding this comment.
Handle zero-valued targets before deriving relative scales
Either reject zero here as requiring strictly positive targets or define a usable scale for it. This check explicitly accepts zero, and the public documentation likewise permits non-negative targets, but calibrate later passes abs(b) / factor to leximin_weights, whose positive-scale validation rejects every zero target. Thus a represented category with a zero total or share always fails despite passing this front-end validation, even when bounds=(0, ...) would make the target reachable.
Useful? React with 👍 / 👎.
Release 0.5.0. Breaking changes are listed in CHANGELOG.md.
calibrate(df, targets, ...),calibrate_replicates,CalibrationReport. The report names the worst-missed targets and says whether other targets or the weight bounds are to blame.tests/test_theory.py.Local checks:
make ci(165 passed, 98.8% coverage),make docs, preen, twine; wheel install plus README example on Python 3.14; tests on numpy 1.26.0 / scipy 1.12.0 / pandas 2.1.1.🤖 Generated with Claude Code