Intro to R
This is a tutorial activity that will introduce you to the basics of R programming. It goes over running code and storing results, working with vectors, building and pulling pieces out of data frames, how folders and file paths work, and how to read a data file into R. It finishes with links for installing R on your own computer.
Every code box on this page runs, including the examples — change them and re-run them if you want to see what happens. Nothing to install — it all runs here in your browser. Click Run to run whatever is in the box, and Check when you think you’ve got it right. If you get stuck, Show hint will point you in the right direction. Your progress saves on this device, so you can stop partway and come back.
You should expect to spend around 10 minutes on this activity.
1. Running code
R is a calculator that remembers things. You type an expression, R evaluates it, and prints the answer. Try running the code below — it’s already complete, so just press Run.
Press Run. The first one takes a few seconds to get going.
The part that makes R useful is storing results. The arrow <- assigns a value to a name, and from then on that name stands in for the value.
Names are case sensitive, can’t start with a number, and shouldn’t contain spaces. mean_age works; Mean Age doesn’t.
Replace the blank so that answer holds the number 42.
The blank is just the number itself: answer <- 42.
Most of the work in R happens through functions. A function has a name, and you give it input in parentheses. sqrt(9) returns 3, round(3.14159, 2) returns 3.14.
Use a function to take the square root of 64 and store it in root.
Square root is sqrt.
2. Vectors
A single number isn’t much use when you have 200 participants. R’s basic container is the vector, which holds several values of the same type. You build one with c(), for “combine”.
Functions like mean(), sd(), min() and max() take a whole vector and return one number. This is the part of R that makes it feel different from a spreadsheet: you describe what you want, not which cells to click.
Build a vector called scores holding 7, 9, 4 and 10, then take its mean.
The combine function is c, so the line starts scores <- c(.
3. Data frames
Real data has several variables at once, so vectors get stacked side by side into a data frame: rows are participants, columns are variables. This is what a dataset looks like in R.
You pull out one column with a dollar sign. If participants is the data frame and age is a column, then participants$age is the vector of ages, and you can hand that straight to mean().
Fill in the column name so mean_age is the average age.
You want the age column, so participants$age.
Pulling pieces out with brackets
The dollar sign hands you a whole column. Square brackets let you ask for particular rows and particular columns, and the order inside the brackets is always rows first, then columns:
Leaving a slot empty means “give me all of them”. The comma does the work here, so participants[2, ] and participants[, 2] ask for completely different things — the second row versus the second column.
If you can remember RC Cola, you can remember the order: Rows, then Columns.
Fill in the blank so third_age is the age of the third participant.
RC Cola: rows first. You want row 3, so participants[3, "age"].
Useful first moves with any new data frame: nrow() for the number of rows, names() for the column names, head() for the first few rows, and summary() for a quick description of every column.
4. Folders and file paths
This is where people usually get stuck, and it has nothing to do with statistics.
R is always sitting in some folder, called the working directory. When you ask it to open a file, it looks starting from there. getwd() tells you where you are.
An absolute path spells out the location from the root of your computer, like /Users/you/Documents/lab-study/data/responses.csv. It works on your machine and nobody else’s. A relative path starts from the working directory instead — data/responses.csv — and keeps working when you send the folder to me, which is why we use relative paths for everything.
So we keep each study in one self-contained folder:
Open lab-study.Rproj and R’s working directory becomes lab-study/, so data/responses.csv points at the right file on any computer. file.path() joins the pieces with the right separator, which is how you avoid the slash-versus-backslash problem between Mac and Windows.
Fill in the folder name so path is "data/responses.csv".
Look at the tree above: the file sits in the data folder.
5. Reading in data
Once the path is right, reading a CSV is one line. read.csv() takes a path and hands back a data frame.
There’s a real file waiting at data/responses.csv in this browser session: eight participants, with id, age, condition and score columns.
Read the file in, then look at the first few rows.
This one's already complete --- just press Run, then Check.
head() prints the column names along the top, which is usually how you find out what you’re actually working with. Now pull one of those columns out and summarise it.
Fill in the column name so avg_age is the average age of the eight participants.
Run the chunk above to see the column names. The one you want is age.
If a read.csv() call fails, the problem is almost always the path rather than the file. Check getwd() first, then check the spelling.
6. Getting R on your own computer
The browser version is a teaching tool. For actual work you want both of these, installed in this order:
- R — the language itself.
- RStudio Desktop — the editor most people use to write R. Free, and it expects R to already be there.
Then make a project rather than a loose script: in RStudio, File → New Project → New Directory, and lay it out like the tree in section 4.
If you want to keep going before we talk, R for Data Science is free online and is where I’d start.
Done?
Finish all 9 steps to unlock your completion record.
Intro to R — completed