An R markdown template is available for this workshop. Copy and paste its content into a new R markdown file in RStudio.
This exercise is concerned with the relation between population size
and foraging area for seabird colonies. Data are available on 22
black-legged kittiwake (a northern gull) colonies on Scotland’s Shetland
and Orkney Islands. The variables recorded are the colony name;
Area, in km2; and Population, the
number of breeding pairs. (The data source is Cairns, D. K. (1988). “The
regulation of seabird colony size: a hinterland model.” The American
Naturalist, 134, 141-146.)
You should be able to read the data directly in R, and stored as the
data frame Kittiwake, by
Kittiwake <- read.csv(file="https://r-resources.massey.ac.nz/data/161251/kittiwake.csv", header=TRUE, row.names=1)
If you Download kittiwake.csv then save it on your computer, use a command like:
## Kittiwake <- read.csv(file="<file path>/kittiwake.csv", header=TRUE, row.names=1)
where you should replace <file path> with the
appropriate address corresponding to the location where you stored the
file on your computer.
head(Kittiwake)
tail(Kittiwake)
Population against
Area using:library(ggplot2)
Kittiwake.scatter <- ggplot(Kittiwake, aes(x=Area, y=Population)) + geom_point()
Kittiwake.scatter
You should see that the scatterplot suggests a non-linear
relationship between the variables, with the curve increasing rapidly
for values of Area. We will therefore need to transform the
data if we are to apply a simple linear regression model.
lPop bylibrary(tidyverse)
Kittiwake |> mutate(lPop = log(Population)) -> Kittiwake
glimpse(Kittiwake)
Now produce a scatterplot of the log of Population
against Area by modifying the ggplot() command
used previously. Store it using the name
Kittiwake.scatter2.
Kittiwake.lm <- lm(lPop ~ 1 + Area, data=Kittiwake)
coef(Kittiwake.lm)
The fitted model is \(\mathbb{E}[lPop] = 6.0611805 + 9.1028218\times 10^{-4} Area\)
Kittiwake.lm1 <- lm(log(Population) ~ Area, data=Kittiwake)
and compare its coefficients.
Kittiwake.scatter2 + geom_smooth(method="lm", se=FALSE)
Note that we “added” the line fitted by lm() on the fly
here. The geom_smooth() didn’t use the model we fitted, but
fitted its own one — the same one as it happens.
\(\log{y} = a + bx\) (same form as our model) is equivalent to y = ea ebx where e is the exponential function that is the inverse of the natural logarithm. This is often re-written as y = Aebx where A=ea.
This means our fitted line is equivalent to \[ \mathbb{E}[\mbox{Population}] = 428.9 e^{0.00091\mbox{Area}} \]
You should compare your work with the solutions for this workshop.