Skip to content

Populace: federal SALT uses PolicyEngine's imputed withholding, not the state liability (revive #1060) #1227

Description

@MaxGhenis

Summary

#1058 found that the emulator's federal SALT deduction uses PolicyEngine's imputed state withholding, not the state income-tax liability that taxsimtest deducts. Draft PR #1060 fixes this, but it was left as a draft because the fix was inert on the eCPS, where mortgage = proptax = 0, so nobody itemizes. #1058 was closed after @feenberg found the two original records matched.

The Populace inputs (#1216) populate proptax, mortgage and otheritem, so 10,011 records itemize in 2022. On them, a different itemized total (v17) is the most common symptom of federal disagreement: 3,844 of the 6,560 TAXSIM itemizers that differ by more than $15 have one, and PolicyEngine's SALT is higher in 3,582 of those. Switching SALT to PolicyEngine's own liability brings 1,071 more TAXSIM itemizers within $15, more than any other single change we tested (adding back Additional Medicare Tax alone: 767).

A counterfactual on the 10,820 records where either engine itemizes, applying #1060's approach (withholding set to PolicyEngine's own liability), raises federal agreement within $15 among TAXSIM itemizers from 34.5% to 45.2%. After netting the separate Additional Medicare Tax and NIIT differences, it goes from 55.2% to 66.1%, and the summed absolute gap falls by a third.

Mechanism (PE-US 2.6.17, identical on policyengine-us main c8aae08d1)

  • variables/gov/irs/income/taxable_income/deductions/itemizing/state_and_local_sales_or_income_tax.py:11-17: max(state_withheld_income_tax + local_income_tax, state_sales_tax + local_sales_tax).
  • .../itemizing/salt_deduction.py:15-23: min(salt_cap, salt).
  • variables/gov/states/tax/income/state_withheld_income_tax.py:10: sums gov.states.household.state_withheld_income_tax, a list of per-state withholding imputations. For example:
    • states/me/tax/income/me_withheld_income_tax.py:12-19: each person's AGI minus the federal single standard deduction, run through Maine's single schedule. It ignores credits, itemizing, exclusions and joint brackets.
    • states/ma/tax/income/ma_withheld_income_tax.py:12-19: each person's AGI minus the MA single exemption, times the Part B rate.
    • New Hampshire has no entry, so withholding is 0 even though the emulator computes NH interest-and-dividends tax.
  • Check on the 2022 data: PE v17 = salt_deduction + mortgage + otheritem for 100% of TAXSIM itemizers, and salt_deduction = min(10,000, max(withheld + local, sales + local sales) + real_estate_taxes) for 99.7%.
  • Among the 3,844 v17-differing records, withholding exceeds PolicyEngine's own liability in 76.6%, with a median ratio of 1.35.

Examples (2022; PolicyEngine intermediate variables computed with the benchmark configuration):

taxsimid state PE state_withheld_income_tax PE state_income_tax TAXSIM siitax TAXSIM SALT (v17 − mortgage − otheritem)
177 ME 5,519.92 1,436.99 586.99 730.56
551 NH 0.00 544.08 544.08 544.08
2449 MA 3,839.15 1,836.14 1,578.51 1,578.51

How much is withholding, and how much is something else

The table shows the share of 2022 TAXSIM itemizers whose SALT equals TAXSIM's implied SALT within $1, under alternative PolicyEngine definitions:

PolicyEngine SALT definition all 10,011 TAXSIM itemizers 3,844 big-gap, v17-differing median |error| in that set
current (imputed withholding) 53.7% 0% $824
PolicyEngine's liability (#1060) 61.2% 19.4% $251
min(10,000, max(TAXSIM siitax, PE state_sales_tax) + proptax) 76.2% 54.4% $0.1

So moving to the liability is the main step. What remains is mostly:

  1. the two engines' state liabilities differing. These are state-model questions, per state, and they flow into federal tax through SALT.
  2. TAXSIM's sales-tax alternative, which is "taken from regressions based on IRS publication 600 ... For tax years after 2016, the 2016 coefficients are used" (taxsimtest docs). PolicyEngine uses the IRS optional table instead. That table is itself misaligned in 30 states (Optional state sales tax table: 30 states hold a neighbouring state's values (column shift), and 2022 uses the 2023 table policyengine-us#9595).

fiitax counterfactual (#1060's approach)

The run sets state_withheld_income_tax := state_income_tax (clipped at 0) before any other calculation, on 2022 records where either engine itemizes. state_income_tax comes from a first pass with the same configuration. This follows #1060's rule but isn't #1060's code. The script is in the collapsed block below.

group n within $15, now → counterfactual after netting addmed & ΔNIIT Σ|gap| (netted)
TAXSIM itemizers 10,011 34.5% → 45.2% 55.2% → 66.1% $1,360,905 → $903,427
PE-only itemizers 809 72.1% → 82.6% 72.1% → 82.6% $81,925 → $59,033

Agreement within 1% of AGI barely moves (93.1% → 93.3%). The dollar errors are small relative to income, but there are many of them.

What the law counts (for the convention decision)

2022 Instructions for Schedule A, line 5a, "State and Local Income Taxes": "If you don't elect to deduct general sales taxes, include on line 5a the state and local income taxes listed next. • State and local income taxes withheld from your salary during 2022. ... • State and local income taxes paid in 2022 for a prior year, such as taxes paid with your 2021 state or local income tax return. ... • State and local estimated tax payments made during 2022, ..."

Neither engine's measure is the statutory one: that is cash paid during the year. TAXSIM uses the current-year liability, which is the natural steady-state proxy. PolicyEngine's per-person withholding imputation is a rougher proxy: it ignores credits and joint brackets, and gives zero for taxes that aren't withheld, such as NH's. For emulating TAXSIM, the liability is the target.

Proposal

Revive #1060 for the Populace benchmark. Its "inert on the eCPS" rationale no longer holds. It brings about 1,070 more TAXSIM itemizers within $15 in 2022 (about 1,090 after netting addmed and NIIT). It doesn't change state output, since #1060 sources the state columns from pass 1. #1058 was closed after the two original records matched. This issue tracks the population-level effect.

Minimal reproducer (Maine joint, 2022): taxsimid 177 in populace_households.csv, with year set to 2022.

  • TAXSIM: fiitax 8,659.62, v17 38,562.07.
  • Emulator: fiitax 8,084.90, v17 43,351.42. That is mortgage 37,831.51 plus withholding 5,519.92.
Self-contained counterfactual script (benchmark environment: policyengine-us 2.6.17, core 3.32.6; run from a checkout of the #1216 branch, which has populace_households.csv)
import numpy as np, pandas as pd
from policyengine_taxsim.runners.policyengine_runner import PolicyEngineRunner, TaxsimMicrosimDataset

YEAR = 2022

def configured_sim(rows):
    """A Microsimulation built exactly as the benchmark builds it (all emulator overrides)."""
    rows = rows.copy(); rows["idtl"] = 5
    runner = PolicyEngineRunner(rows, logs=False, assume_w2_wages=True, disable_salt=False)
    df = runner.input_df
    df["year"] = df["year"].astype(float).astype(int)
    ds = TaxsimMicrosimDataset(df); ds.generate()
    return runner._build_configured_sim(ds, df), ds

inputs = pd.read_csv("populace_households.csv").assign(year=YEAR)
out = []
for start in range(0, len(inputs), 10_000):
    chunk = inputs.iloc[start:start + 10_000].reset_index(drop=True)
    sim1, ds1 = configured_sim(chunk)            # pass 1: the state liability
    liability = sim1.calculate("state_income_tax", YEAR).values.clip(min=0)
    itemizes = sim1.calculate("tax_unit_itemizes", YEAR).values
    base_tax = sim1.calculate("income_tax", YEAR).values
    ds1.cleanup()
    sim2, ds2 = configured_sim(chunk)            # pass 2: SALT from the liability
    sim2.set_input("state_withheld_income_tax", YEAR, liability)
    out.append(pd.DataFrame({
        "taxsimid": chunk.taxsimid,
        "pe_itemizes": itemizes,
        "fiitax_now": base_tax,
        "fiitax_liability_salt": sim2.calculate("income_tax", YEAR).values,
        "niit_liability_salt": sim2.calculate("net_investment_income_tax", YEAR).values,
    }))
    ds2.cleanup()
pd.concat(out).to_csv("cf_liability_2022.csv", index=False)

Compare fiitax_liability_salt against TAXSIM's fiitax from comparison_results_2022.csv in the release. The figures above restrict to records where TAXSIM or PE itemizes (10,820).

Related: #1058, #1060, #971, #997, #1216.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions