Skip to content

distDD: keep every bin in the table when one is dropped - #7

Open
brycewang-stanford wants to merge 1 commit into
pedrohcgs:mainfrom
brycewang-stanford:fix-distdd-dropped-bin
Open

brycewang-stanford wants to merge 1 commit into
pedrohcgs:mainfrom
brycewang-stanford:fix-distdd-dropped-bin

Conversation

@brycewang-stanford

Copy link
Copy Markdown

distDD() stops when one of the bins has a degenerate influence function and is dropped. On mpdta this happens with aggte_type = "dynamic" and balance_e = 1 or 2.

library(didFF)
data(mpdta, package = "did")

distDD(data = mpdta, yname = "lemp", tname = "year", idname = "countyreal",
       gname = "first.treat", control_group = "nevertreated",
       nbins = 6, seed = 0, aggte_type = "dynamic", balance_e = 1)
#> Error in data.frame(level = bin2, test.estimates = point_estimates,  :
#>   arguments imply differing number of rows: 6, 5

The same call without balance_e, or with balance_e = 0, returns six rows, and didFF() with the same arguments runs. I see it with both did 2.3.0 and did 2.5.1.

Cause

didFF() drops bins whose influence function has close to zero variance (keep_IF) and builds Sigmahat from the remaining columns. The distDD branch at the end still pairs bin2 and point_estimates, which hold every bin, with sqrt(diag(Sigmahat)), which holds the kept ones. The didFF branch is not affected, because its table has level and implied_density only.

Change

The table keeps one row per bin, and a dropped bin gets NA as its standard error. In the example above the last bin is the dropped one. Its estimate is exactly zero and its standard error is now NA:

                                            level test.estimates     test.se
1 [1.08926820442608396355,2.65595966233907310183]  -0.0016181230 0.005720855
2 (2.65595966233907310183,4.21330703601003619951]   0.0161272923 0.017398012
3 (4.21330703601003619951,5.77065440968099974128]  -0.0125943905 0.038540929
4 (5.77065440968099974128,7.32800178335196328305]  -0.0008360302 0.038577388
5 (7.32800178335196328305,8.88534915702292593664]  -0.0010787487 0.016308576
6 (8.88534915702292593664,10.4520406149359139647]   0.0000000000          NA

Calls where no bin is dropped return what they returned before.

NA seemed the least surprising choice to me, since the rows stay aligned with the bins. If you would rather drop the row, or report a zero, that is a one-line change and I am glad to adjust.

Testing

tests/test_distDD.R is new. It runs the example with and without balance_e = 1 and checks that both tables have six rows with the same levels and that exactly one standard error is missing in the second. It errors on main and passes with this change (did 2.5.1, R 4.5.2).

The test passes control_group explicitly, so it does not depend on #6.

didFF() drops bins whose influence function has close to zero variance
and builds Sigmahat from the rest. The distDD table still paired all
bins with diag(Sigmahat), so data.frame() stopped with "arguments imply
differing number of rows" (e.g. mpdta, aggte_type = "dynamic",
balance_e = 1).

Report NA as the standard error of a dropped bin.
@brycewang-stanford

Copy link
Copy Markdown
Author

This is independent of #6 and the two can be merged in either order. I found both while checking StatsPAI's distributional DiD against distDD; the background is in my comment on #6.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant