Research questions, research design

Sept 2026

Macartan Humphreys

1 Today

  • Questions
  • Arguments as DAGs
  • Hypotheses
  • Strategies
  • Designs and design diagnosis
  • Workflow

2 Questions

2.1 Types of claims

\[X \rightarrow Y\]

2.2 Types of claims

\[X \rightarrow Y\]

  • Analytic claims: e.g. \(X=1\) implies \(Y=1\)
  • Descriptive claims: e.g. \(Y=1\) when \(X=1\)
  • Causal claims (interventionist): e.g. \(X=1\) causes \(Y=1\)
  • Causal claims (counterfactual): e.g. \(Y=1\) because \(X=1\)

We mostly focus on causal claims. Even claims we think of as descriptive are often causal claims.

2.3 What are \(X\) and \(Y\)?

\[X \rightarrow Y\]

  • Names for \(X\): independent variable, explanatory variable, input, exogeneous variable, cause, driver, right hand side variable

  • Names for \(Y\): dependent variable, output, outcome, endogenous variable, left hand side variable

  • Both implicitly have a location and a timestamp: “\(Y=1 \Leftrightarrow\) The US was a democracy in 2000

2.4 What’s the question?

  • \(X \rightarrow Y\)
  • \(? \rightarrow Y\)
  • \(X \rightarrow ?\)
  • \(? \rightarrow ?\)

And:

  • Does \(X\) affect \(Y\) in general? (effects of causes)
  • Did \(X=1\) cause \(Y=1\) in this case? (causes of effects)

2.5 Other types of variables

  1. mediating variables
  2. conditioning or moderating variables
  3. confounding variables
  4. instrumental variables

2.6 Mediating variables

  • “Oil wealth produces grievances which cause conflict”

2.7 Moderating variables

  • e.g. “The effect of oil wealth on conflict is weaker when institutions are strong”

2.8 Mediation and moderation

  • “Oil wealth produces grievances which cause conflict”
  • “There are many other channels through which oil wealth affects conflict and which exacerbate the effects of grievances”

2.9 Confounding variables

  • “The effect of education on voting behavior is hard to assess because wealth affects both education and voting behavior”

2.10 Instrumental variables

“It’s hard to assess the effect of military service on future earnings because of individual characteristics that might explain both. But date of birth affects the chances of serving and so can be used to recover estimates of service on earnings.”

3 Models: Arguments as DAGs

3.1 An argument:

Here is a complete, albeit barebones (and possibly incorrect), argument:

  • Good institutions (I) cause economic growth (G), except in countries with large stocks of natural resources (N)
  • The reason is that institutions encourage people to invest (V) which spurs growth (this effect does not kick in in natural resource rich countries as people just live off rent)
  • Growth also makes it easier to maintain good institutions, which creates a virtuous cycle
  • Being an ally (A) of the US helps growth, but it can corrupt domestic institutions.
  • Historically, places with climates (C) suitable for colonizers to settle had better institutions. These climatic conditions are otherwise irrelevant for contemporary economic growth.

3.2 Some counterarguments:

  • Places with climates suitable for colonizers benefited from better access to international markets which led to growth.
  • Good soil is also important for growth!
  • Good institutions also make sure that investments yield greater returns and that’s what causes growth

3.3 Questions on Nodes

  • What are the dependent variables?
  • What are the independent variables?
  • What are the mediating variables?
  • What are the conditioning variables?
  • What are the confounding variables?
  • What are the instrumental variables?
  • Graph the relations between the variables.

3.4 Questions on Inference

  • Say I and G are positively correlated. Does this mean that I causes G?

  • Say I and G are negatively correlated. Does this mean that I does not cause G?

  • How might you estimate the effect of I on G?

  • How does C help establish the link between I and G?

  • Where is the theory? Is in equivalent to the graph or is it something else that generates the graph?

  • How might you check if the proposed theory is correct?

  • Which of the counterarguments are strong and why?

3.5 A graph

3.6 Exercises: Dissect these arguments

Four arguments. For each one you should identify the:

  • type of argument (effect of \(X\), cause of \(Y\), effect of \(X\) on \(Y\))
  • unit of analysis
  • dependent variable(s)
  • independent variable(s)
  • mediator(s)
  • possible conditioning variable(s)
  • possible confounder(s)
  • possible identification strategy
  • relevant key agent(s) (actor(s))
  • measurement strategy

3.7 A. Natural resources and conflict

In developing countries that discover natural resources, such as oil, the ruling elite can extract wealth without needing to tax citizens and develop the state apparatus. Because the state does not rely on taxation for government revenue, it does not need to set up accountability structures or extend its reach and citizens do not feel that they have ownership over the state. The state therefore becomes both less democratic and weaker than if it had not discovered the resources.

3.8 B. Democracy and growth

Rich countries are more likely to be democratic for the simple reason that when people become wealthier they refuse to be dictated to by others and they demand a role in government. The marginal effects of income increases are greater for poorer countries because the impacts on eduction are greatest at these levels. You can test this proposition by exploiting natural variation in commodity prices which provide shocks to national income, especially for countries dependent on primary commodity exports.

3.9 C. Factor Endowments and Coalitions

When countries increase trade (imports and exports), the returns to economic factors (such as labor, land and capital) are affected differently. Specifically, the returns to factors that are the most abundant are positive, while the returns to factors that are the most scarce are negative. Therefore, the relative factor endowments of a country will predict what sort of political coalitions will form (eg Land versus Labor + Capital) and which groups will favor free trade policies.

3.10 D. Democratic peace

In democratic states, leaders are accountable for any losses incurred as a result of the wars that they enter into. Two states with democratic leaders are also more likely to share a common set of norms, and to engage in trade with one another. Therefore, two democracies are far less likely to enter into war with one another than a democracy and a non-democracy, or two non-democracies.

4 Hypotheses

4.1 Take home ideas

  • You don’t need them, but stating expectations in terms of hypotheses provides discipline to a research project.

  • Hypotheses are statements about the world that you seek to reject

  • A good hypothesis is simple and falsifiable

  • A \(p\) value is the probability of data like what you see under some particular hypothesis

4.2 Characteristics of good hypotheses

  • They are possibly TRUE or FALSE
  • They are falsifiable
  • They are statements about the world, not your analysis
  • They are simple (not double barreled)
  • They involve clear concepts
  • They are few, and they are motivated
  • They are contested: You will learn something whether the data supports them or rejects them. Most importantly: you are not sure if they are true or false
  • They are numbered, and maybe even named

4.3 Some hypotheses

Consider these:

  • Education is very important
  • Education increases your income
  • Education either increases, decreases, or has no effect on your income
  • Education is good for you because it strengthens your character in very fundamental ways that you could never measure

Now:

  • Just one of these is not a hypothesis. Which one?
  • Just one of these is a good hypothesis. Which one?

4.4 Nulls: A point of confusion

Because of an unusual convention, social scientists often describe hypotheses in terms of what they expect but then test the null hypothesis of no effect

eg:

  • H1: Competition reduces prices
  • H-null: Competition has no effect on prices

5 Identification strategies

5.1 Define the average causal effect

  • \(Y_i(1)\) is the outcome \(i\) would have if \(X\) were 1
  • \(Y_i(0)\) is the outcome \(i\) would have if \(X\) were 0
  • \(Y_i\) is what’s actually observed, given \(X\)
  • \(Y_i(1)- Y_i(0)\) is the unit level treatment effect
  • \(\frac{1}n \sum_{i}\left(Y_i(1)- Y_i(0)\right)\) is the average treatment effect

So: the average treatment effect is just the average of the differences between the what the outcome would be in treatment and what the outcome woud be in control for each unit. It is unfortunately not measurable!

5.2 Differences in means works if you have randomization

  • We want to estimate: \[\frac1n\sum_i(Y_i(1) - Y_i(0))\]

  • We estimate using: \[\frac1{n_t}\sum_{i \in T}Y_i - \frac1{n_c}\sum_{i \in C}Y_i\]

  • This works because, with randomization \(\frac1{n_t}\sum_{i \in T}Y_i = \frac1n\sum_i((Y_i(1))\) in expectation – that is, on average the sample average is the population average. Similarly \(\frac1{n_c}\sum_{i \in C}Y_i = \frac1n\sum_i((Y_i(0))\) in expectation.

  • “The difference in averages is the same as the the average of differences”

5.3 Otherwise there are threats to inference

Difficulties once assignment is related to potential outcomes.

Here X might be related to Y even though it does not cause Y

5.4 In fact in the absence of randomization a model is required

  • Randomization is not required for causal inference.

  • But without it you need some alternative argument for why your estimates from the treatment group capture what would occur in the control group if they were treated (and vice versa)

  • What’s more you will need a model:

    • Some variables might have to be taken into account in order to ensure no confounding
    • Some variables might have to not be taken into account in order to ensure no confounding
    • Sometimes these might be the same variables!

5.5 Alternatives

Randomization provides a very useful benchmark. Other strategies seek to approximate the magic of randomization:

  • Adjustments. Controlling: Regression, Matching and Weighting
  • Instrumental variables or Natural experiments—seeks a shock that approximates randomization.
  • Difference in differences—assumes that once you account for common time trends cases are as-if randomized
  • Regression discontinuity—assumes that cases are as-if randomized around a threshold (there are also motivations for RDD that do not assume as-if randomization)
  • Synthetic Matching and other model based approaches

5.6 Adjustment methods: Intuition

Key idea is to figure out effects conditional on the values others nodes my take.

Our problem:

5.7 Adjustment methods: Intuition

Key idea is to figure out effects conditional on the values others nodes my take.

Our solution:

\[X \rightarrow Y \>\>\> \> (\text{given }W = 0)\] \[X \rightarrow Y \>\>\> \> (\text{given }W = 1)\] We can estimate effects within similar sets and then average the results (weighting by the size of the sets)

5.8 RDD Intuition:

e.g. compare units in which the margin of victory was 1 vote for democrats against those for which it was -1. We expect that these are approximately identical, on average, is all regards and get an estimate of the effect of victory on some outcome at the threshold

5.9 Diff-in-diff Intuition:

  • examine the difference between treated and control in the before-after difference
  • OK under the assumption that the change in the control group is the same as what the change would have been in the treatment group absent treatment

6 Design declaration and diagnosis

…pulls a lot of these elements of a design together

book: https://book.declaredesign.org/

6.1 The MIDA Framework

Four elements of any research design:

  • Model: set of models of what causes what and how
  • Inquiry: a question stated in terms of the model
  • Data strategy: the set of procedures we use to gather information from the world (sampling, assignment, measurement)
  • Answer strategy: how we summarize the data produced by the data strategy

6.2 Four elements of any research design

6.3 Examples of MIDA elements

  • M: DAGs, game theoretic models
  • I: ATEs, CATEs, COEs, models
  • D: Sampling schemes, assignment schemes, text analysis, interview
  • A: Experiment, observational, quantitative, qualitative:
    • Conditioning on observables
    • Difference in differences
    • RDD
    • Instrumental variables

6.4 Declaration, Diagnosis, Redesign cycle

  • Declaration: Telling the computer what M, I, D, and A are.

  • Diagnosis: Estimating “diagnosands” like power, bias, rmse, error rates, ethical harm, amount learned.

  • Redesign : Fine-tuning features of the data and answer strategies to understand how they change the diagnosands

  • Different sample sizes

  • Different randomization procedures

  • Different estimation strategies

  • Implementation: effort into compliance versus more effort into sample size

7 Topics: Workflow

7.1 Classic paper structure

https://macartan.github.io/teaching/how-to-write

Classic structure

  1. Motivation
  2. Theory
  3. Strategy (perhaps descriptives here)
  4. Main results
  5. Discussion / deepening
  6. Conclusion

7.2 Classic paper stages

  1. Come up with a question
  2. Come up with an answer strategy, find a data source
  3. Develop the design
  4. Present the design, modify as needed
  5. Register the design
  6. Get ethics approval if needed
  7. Implement data strategy
  8. Implement answer strategy; generate replication material in parallel
  9. Generate slides and present to colleagues
  10. Submit

7.3 Principles

  • Always work from a folder where your work automatically backs up.

    • Dropbox, Drive, many others
  • Have analysis files integrated with writing files

    • qmd fabulous for this
    • Set it up so that you can quickly do reality checks on data and analysis
  • Be able to replicate all data work and analysis with one click

  • Outsource formatting.

    • e.g. tex, .qmd automatically format. If you use Word, use their “Styles”
    • Bibliographies: use bibtex or similar. Keep a file and reference like this @putnam2000bowling (in qmd) which produces Putnam (2000) and handles the formatting. Other tools work similarly. Don’t do this by hand.

7.4 Folders and files

  • Have all files: writing, files, data files, additional analysis files or images, etc in a single directory with relative references

    • Number your folders.
    • Have few files in each folder.
    • Compile regularly and have a readable compiled file beside your work file
    • Keep your main document clean
    • Have an archive folder 0_archive and backup old copies regularly (not so important if you have good versioning; I often label backup files with date: 20201005_paper.qmd)

7.5 Examples

See examples in sample_project or student folders

7.6 AI

Do not outsource your thinking

Follow general TCD policy

  • Don’t let AI do any drafting of your text; but you can ask a good LLM if your text is clear or for advice on how to improve (but be wary of any advice you get).
  • Don’t let AI select analysis strategies for you; you might ask it to check your code and look for bugs and suggest improvements; but you have to decide how to act on any advice you get.
  • Don’t rely on AI for literature searches or summaries; it can provide suggestions for entryways but any source you draw on you should consult yourself and evaluate for yourself.

7.7 AI

You will be questioned on your analyses, your knowledge of literatures, your results and how to interpret them; you will need to understand them thoroughly and be able to defend your choices.

Most important principle is that you can vouch for everything that you submit. No references to articles you have not read. No arguments or methods you do not understand. No text you have not reviewed.

7.8 Tasks

  • Keep a to-do list
  • If you use github it is great to use “issues” to keep track of to dos. I always have a todo memo in each project folder.
  • Try to complete well defined tasks in one sitting. Work in chunks.
  • If you have a repetitive task there is probably a way to automate it: ask for help
  • If you have a conceptually hard task that you are not making progress on, stop, move away, and try it from a whole new angle
  • Order your tasks: figure out whether you work better linearly or doing parallel work.
  • Cross off tasks when done
  • Don’t be afraid to discard work
  • Be ambitious but don’t let the best be the enemy of the good.

8 More

Design:

Blair et al. (2023), King et al. (2021)

Writing:

Strunk Jr and White (2007), Lipson (2018)

9 References

Blair, Graeme, Alexander Coppock, and Macartan Humphreys. 2023. Research Design in the Social Sciences: Declaration, Diagnosis, and Redesign. Princeton University Press.
King, Gary, Robert O Keohane, and Sidney Verba. 2021. Designing Social Inquiry: Scientific Inference in Qualitative Research. Princeton university press.
Lipson, Charles. 2018. How to Write a BA Thesis: A Practical Guide from Your First Ideas to Your Finished Paper. University of Chicago Press.
Putnam, Robert D. 2000. “Bowling Alone: America’s Declining Social Capital.” In Culture and Politics. Springer.
Strunk Jr, William, and Elwyn Brooks White. 2007. The Elements of Style Illustrated. Penguin.