Lab 2: Wrangling

How labs work

You are encouraged to work together and get started in class—talk through the problems with the people around you. You will finish the lab on your own time and submit your own work.

AI use: You may use AI to ask questions or debug, but not to write your code for you—the point is to think it through yourself. At the end of the lab, add a one-line note saying whether and how you used AI.

Overview

In Lab 1 we handed you clean data. Now you will build a tidy data set yourself—the skill that makes every later visualization possible. You will pull a democracy measure from V-Dem and a development or empowerment indicator from the World Bank, wrangle them, merge them into one data frame, and finish with a chart comparing regions.

Step 1: Download your data (20 pts)

a) Use vdem from vdemdata to download the V-Dem dataset. Use filter() to keep the years from 1990 onward, then use select() to keep and rename the country name, the V-Dem country ID, the year, a democracy measure of your choice and the region variable (e_regionpol_6C). Recode the region numbers into names with mutate() and case_match(). glimpse() the result.

b) Use wb_data() from wbstats to download GDP per capita (or another development or women’s-empowerment indicator) for the same years. Give your indicator a readable name, drop iso2c, and rename date to year so it matches V-Dem. glimpse() the result.

Step 2: Add country codes (15 pts)

V-Dem and the World Bank use different country codes, so we need a common one before we can merge. Use mutate() and countrycode() to add an iso3c variable to your V-Dem data frame, converting from the "vdem" codes to the "wb" codes. glimpse() the result to confirm the new column is there.

Step 3: Merge (20 pts)

Use left_join() to merge your V-Dem and World Bank data frames by iso3c and year. Make sure you end up with just one country name column. glimpse() the result and check for excessive NAs.

Step 4: Summarize by region (20 pts)

a) Starting from your merged data frame, use group_by() and summarize() to calculate the mean (or median) of one of your variables—your democracy measure or your World Bank indicator—for each region. You can use all years or filter() for a single year first. Save the result as a new object.

b) Use arrange() to sort the regions from highest to lowest. Which region comes out on top? Which is at the bottom?

Step 5: Visualize (25 pts)

a) Make a column chart with geom_col() showing your summarized variable by region. Sort the bars from highest to lowest, choose a fill color and a theme, and add axis labels, a title and a caption naming your source.

b) Below this line, write a sentence or two describing what the chart shows. Is this what you expected?


AI statement: (one line on whether and how you used AI)


Hints

Only look at these if you’re stuck!

Hint 1 — Recoding the region variable:

e_regionpol_6C comes as numbers from 1 to 6. Use case_match() inside mutate() to turn them into names:

mutate(
  region = case_match(region,
                      1 ~ "Eastern Europe",
                      2 ~ "Latin America",
                      3 ~ "Middle East",
                      4 ~ "Africa",
                      5 ~ "The West",
                      6 ~ "Asia")
)

Hint 2 — Naming your World Bank indicator:

If you pass wb_data() a named vector, the column comes back with your name instead of the indicator code:

indicators <- c("gdp_pc" = "NY.GDP.PCAP.CD")

wb_data(indicators, start_date = 1990, end_date = 2023)

Hint 3 — Do we need to reshape?

No. Both V-Dem and wb_data() already give you one row per country-year, so the two data frames have the same shape. You only need to make sure the columns you join on have the same names (year, not date) and the same codes (that is what countrycode() is for).

Hint 4 — The countrycode() warning:

You will probably see a warning like Some values were not matched unambiguously. That is OK. V-Dem includes a few historical units (like Zanzibar or Somaliland) that have no World Bank code, so they get an NA for iso3c and will not find a match in the merge.

Hint 5 — Two country columns after the merge:

Both data frames have a country column, so left_join() gives you country.x and country.y. Drop the World Bank version with select(-country) before you join, or drop and rename the extra column after.

Hint 6 — NAs in your summary:

If your summary comes back as NA, it is because some country-years are missing data. Add na.rm = TRUE:

summarize(gdp_pc = mean(gdp_pc, na.rm = TRUE))

Hint 7 — Sorting the bars:

arrange() sorts your table but not your chart—ggplot() will still put the regions in alphabetical order. Use fct_reorder() inside aes() to sort the bars (the - puts the highest value first):

ggplot(region_summary, aes(x = fct_reorder(region, -gdp_pc), y = gdp_pc)) +
  geom_col(fill = "steelblue")

Hint 8 — Formatting the y-axis:

If you chart GDP per capita, scale_y_continuous(labels = label_dollar()) from the scales package will format the axis as dollars. For a percentage, use label_percent(scale = 1).