Computer Laboratory Exercise 2B

An R markdown template is available for this workshop. Copy and paste its content into a new R markdown file in RStudio.

In this exercise you will:

  • add the fitted line from a model to a scatter plot
  • take the model fitted on the transformed scale back to the original scale and present it on a graph

Continuing the Analysis of Kittiwake Colony Data

This exercise is concerned with the relation between population size and foraging area for seabird colonies. Data are available on 22 black-legged kittiwake (a northern gull) colonies on Scotland’s Shetland and Orkney Islands. The variables recorded are the colony name; Area, in km2; and Population, the number of breeding pairs. (The data source is Cairns, D. K. (1988). “The regulation of seabird colony size: a hinterland model.” The American Naturalist, 134, 141-146.)

You should be able to read the data directly in R, and stored as the data frame Kittiwake, by

Kittiwake <- read.csv(file="https://r-resources.massey.ac.nz/data/161251/kittiwake.csv", header=TRUE, row.names=1)

If you Download kittiwake.csv then save it on your computer, use a command like:

## Kittiwake <- read.csv(file="<file path>/kittiwake.csv", header=TRUE, row.names=1)

where you should replace <file path> with the appropriate address corresponding to the location where you stored the file on your computer.

  1. Look at the data frame by simply typing its name, but also try:
head(Kittiwake)
tail(Kittiwake)
  1. Reproduce a scatterplot of Population against Area using:
library(ggplot2)
Kittiwake.scatter <- ggplot(Kittiwake, aes(x=Area, y=Population)) + geom_point()
Kittiwake.scatter

You should see that the scatterplot suggests a non-linear relationship between the variables, with the curve increasing rapidly for values of Area. We will therefore need to transform the data if we are to apply a simple linear regression model.

  1. Create a new variable lPop by
library(tidyverse)
Kittiwake |> mutate(lPop = log(Population)) -> Kittiwake
glimpse(Kittiwake)

Now produce a scatterplot of the log of Population against Area by modifying the ggplot() command used previously. Store it using the name Kittiwake.scatter2.

  1. Reproduce the simple linear regression model to the transformed data by
Kittiwake.lm <- lm(lPop ~ 1 + Area, data=Kittiwake)
  1. Extract the parameter estimates for the fitted model equation using:
coef(Kittiwake.lm)

The fitted model is \(\mathbb{E}[lPop] = 6.0611805 + 9.1028218\times 10^{-4} Area\)

  1. Look at the model specified slightly differently, using:
Kittiwake.lm1 <- lm(log(Population) ~ Area, data=Kittiwake)

and compare its coefficients.

  1. We can update our graph using:
Kittiwake.scatter2 + geom_smooth(method="lm", se=FALSE)

Note that we “added” the line fitted by lm() on the fly here. The geom_smooth() didn’t use the model we fitted, but fitted its own one — the same one as it happens.

  1. No one thinks of populations on a log scale.

\(\log{y} = a + bx\) (same form as our model) is equivalent to y = ea ebx where e is the exponential function that is the inverse of the natural logarithm. This is often re-written as y = Aebx where A=ea.

This means our fitted line is equivalent to \[ \mathbb{E}[\mbox{Population}] = 428.9 e^{0.00091\mbox{Area}} \]

Solutions

You should compare your work with the solutions for this workshop.