Climate-smart agriculture in Zambia
Endogenous Switching Regression
Nobody randomised who adopts. So you have to build the world in which they did not.
Nobody randomised who adopts. Farmers chose, so adopters and non-adopters differ in ways that go beyond the practice itself, and comparing them directly measures the practice plus the kind of household that takes it up.
A switching regression models the decision to adopt and the outcome together, then constructs the counterfactual: what an adopting household would have earned had it not adopted, and what a non-adopting household would have earned had it adopted. Those two invented worlds are the comparison.
It is transparent and it is interpretable in ordinary economic terms, which is why it remains the traditional choice in this literature.
Getting the uncertainty right
There is a trap in two-stage models like this one. The second stage uses a term produced by the first, and if you then report standard errors as though that term had been handed to you as a known quantity, you understate the uncertainty and the results look sharper than they are.
The repair is to resample the whole procedure, both stages together, many times over, and read the spread of the answers. It is slower and it is more honest, and it can turn a table of confident-looking results into a much quieter one.
What it cannot do
You have to write down the shape of every relationship in advance. Write down the wrong shape and the estimates are wrong, and the model will not tell you. That is the specific weakness double machine learning was brought in to cover.
The equation
choice: Ij* = zγj + ηj outcome in regime j: yj = xβj + σjλj + uj
- Ij*
- the unobserved appeal of bundle j to this household. It picks the highest
- yj
- income or yield, with its own equation for each bundle rather than one shared line
- λj
- the selectivity term, carried over from the choice equation
- σj
- how strongly the choice and the outcome are linked
λ is the part that does the work. It is the model’s way of admitting that whatever drew a household towards a bundle may also be raising its income, and of subtracting that before reporting an effect.
Further reading
- Lee, L.-F. (1983). "Generalized Econometric Models with Selectivity", Econometrica. Read it.The selection correction this model uses.
- Bourguignon, F., Fournier, M. and Gurgand, M. (2007). "Selection Bias Corrections Based on the Multinomial Logit Model", Journal of Economic Surveys. Read it.How the correction behaves when the choice is among several options rather than two.
From the Zambia climate-smart agriculture research, built on the Water and Soil Accelerator household survey, which was funded by USAID. The thesis is under examination and the three papers drawn from it are under anonymous peer review, so there is nothing to link to yet.