---
title: Massey University --- School of Mathematical and Computational Sciences  
subtitle: 161.251 Regression Modelling
author: "Staff member responsible for this workshop: Jonathan Godfrey a.j.godfrey@massey.ac.nz"
date: "last updated 11 July 2025"
output:
    html_document
---


```{r setup, include=FALSE}
knitr::opts_chunk$set(comment="", dev="png")
library(tidyverse)
```


N.B. Please note the following:

- Read the full workshop before using this template. External images are not included in your template.
- The above chunk is setting some global options for the rest of this document. It isn't strictly necessary.
- The template worked on my computer, but I have access to all external files, including the data and images.
- Before you proceed, test this template works for you.


# Computer Laboratory  Exercise 2B 

An [R markdown template](https://R-Resources.massey.ac.nz/161251/labs/Lab2B.Rmd) is available for this workshop. Copy and paste its content into a new R markdown file in RStudio.


## In this exercise you will:

- add the fitted line from a model to a scatter plot
- take the model fitted on the transformed scale back to the original scale and present it on a graph

## Continuing the Analysis of Kittiwake Colony Data


This exercise is concerned with the relation between population size and
foraging area for seabird colonies. Data are available on 22
black-legged kittiwake (a northern gull) colonies on Scotland's Shetland
and Orkney Islands. The variables recorded are the colony name;
`Area`, in km^2^; and `Population`, the number
of breeding pairs. (The data source is Cairns, D. K. (1988). "The regulation of seabird colony size: a hinterland model." *The American Naturalist*, **134**, 141-146.)


You should be able to read the data directly in R, and stored as the
data frame `Kittiwake`,
    by

```{r}
Kittiwake <- read.csv(file="https://r-resources.massey.ac.nz/data/161251/kittiwake.csv", header=TRUE, row.names=1)
```


If you  `<a href="data:text/csv;base64,IkNvbG9ueSIsIkFyZWEiLCJQb3B1bGF0aW9uIg0KIlcuIFVuc3QiLDIwNy43LDMxMQ0KIkhlcm1hbmVzcyIsMTU3MCwzODcyDQoiTi5FLiBVbnN0IiwxNTg4LDQ5NQ0KIlcuIFllbGwiLDEyNS43LDEzNA0KIkJ1cmF2b2UiLDM1My40LDQ4NQ0KIkZldGxhciIsOTMxLDM3Mg0KIk91dCBTa2VycmllcyIsMTYxNiwyODQNCiJOb3NzIiwxMzE3LDEwNzY3DQoiTW91c3NhIiw2MTQsMTk3NQ0KIkRhbHNldHRlciIsNjAuMTIsOTcwDQoiU3VtYnVyZ2giLDEyNzMsMzI0Mw0KIkZpdGZ1bCBIZWFkIiw1OTUuNyw1MDANCiJTdCBOaW5pYW5zIiwxMDUuNiwyNTANCiJTLiBIYXZyYSIsMjQxLjksOTI1DQoiUmVhd2ljayIsMTExLjEsOTcwDQoiVmFpbGEiLDMwMi40LDI3OA0KIlBhcGEgU3RvdXIiLDgwOC45LDEwMzYNCiJGb3VsYSIsMjkyNyw1NTcwDQoiRXNoYW5lc3MiLDEwNjksMjQzMA0KIlV5ZWEiLDg5OC4yLDczMQ0KIkdydW5leSIsNTY0LjgsMTM2NA0KIkZhaXIgSXNsZSIsMzk1NywxNzAwMA0K" download="kittiwake.csv">Download kittiwake.csv</a>`{r} then save it on your
computer, use a command like: 


``` r
## Kittiwake <- read.csv(file="<file path>/kittiwake.csv", header=TRUE, row.names=1)
```
where you should replace `<file path>` with the appropriate address corresponding to the location where you stored the file on your
computer.

1. Look at the data frame by simply typing its name, but also try:

``` r
head(Kittiwake)
tail(Kittiwake)
```

1. Reproduce a scatterplot of `Population`
against `Area` using:

``` r
library(ggplot2)
Kittiwake.scatter <- ggplot(Kittiwake, aes(x=Area, y=Population)) + geom_point()
Kittiwake.scatter
```

You should see that the scatterplot suggests a non-linear relationship
between the variables, with the curve increasing rapidly for values of
`Area`. We will therefore need to transform the data if we
are to apply a simple linear regression model.

2. Create a new variable `lPop` by


``` r
library(tidyverse)
Kittiwake |> mutate(lPop = log(Population)) -> Kittiwake
glimpse(Kittiwake)
```

Now produce a scatterplot of the log of `Population` against
`Area` by modifying the `ggplot()` command used previously. Store it using the name `Kittiwake.scatter2`.


3. Reproduce the simple linear regression model to the transformed data by

``` r
Kittiwake.lm <- lm(lPop ~ 1 + Area, data=Kittiwake)
```


4. Extract the parameter estimates for the fitted model equation using:


``` r
coef(Kittiwake.lm)
```

The fitted model is $\mathbb{E}[lPop] = 6.0611805 + 9.1028218\times 10^{-4} Area$


5. Look at the model specified slightly differently, using:


``` r
Kittiwake.lm1 <- lm(log(Population) ~ Area, data=Kittiwake)
```

and compare its coefficients.


6. We can update our graph using:


``` r
Kittiwake.scatter2 + geom_smooth(method="lm", se=FALSE)
```

Note that we "added" the line fitted by `lm()` on the fly here. The `geom_smooth()` didn't use the model we fitted, but fitted its own one --- the same one as it happens.


7. No one thinks of populations on a log scale.


$\log{y} = a + bx$ (same form as our model) is equivalent to *y = e^a^ e^bx^* where *e* is the exponential function that is the inverse of the natural logarithm. This is often re-written as *y = Ae^bx^* where *A=e^a^*.

This means our fitted line is equivalent to $$
\mathbb{E}[\mbox{Population}] = 428.9   e^{0.00091\mbox{Area}}
$$


## Solutions

You should compare your work with the [solutions for this workshop.](https://R-Resources.massey.ac.nz/161251/labs/Lab2B-sols.html)
