R is a language built for statistics and working with data. Numbers come in vectors
rather than one at a time, so a whole column is summed, filtered or plotted in a
single line, and the tidyverse packages add one clear verb for each data job. The
reference below is grouped by what you are trying to do, and the filter box
searches all of it at once. Type dplyr to pull out the data-frame verbs, or
4.5 to see what arrived in recent releases.
Every snippet is checked against R 4.6, the current version, and was run on
R 4.6.1 with dplyr 1.2.1, ggplot2 4.0.3, readr 2.2.0 and tidyr 1.3.2. Anything
newer than R 4.1 says so in the notes column. Names like x, df and user are
placeholders for your own, and the data examples use penguins, which ships with
R from 4.5. Coming from spreadsheets or databases? The
SQL cheat sheet covers the same filtering and
grouping on the database side, and the Python cheat sheet
is grouped the same way as this page.
Searches the task, the command and the third column. Press / from anywhere on the page.
283 commands
Running R
| Task | Command | Notes |
|---|---|---|
| Check which version you have | R --version | Inside R: R.version.string |
| Run a script | Rscript analysis.R | |
| Run one line without a file | Rscript -e "print(2^10)" | |
| Open the interactive console | R | Leave with q(). Answer n to saving the workspace |
| Run a script from inside R | source("analysis.R") | |
| Read the arguments passed to a script | args <- commandArgs(trailingOnly = TRUE) | Rscript analysis.R sales.csv gives args[1] as "sales.csv" |
| Help for a function | ?mean help("mean") | ??regression searches every installed help page |
| Run the examples from a help page | example(mean) | |
| Where relative paths are read from | getwd() | setwd() changes it. Open the project folder instead of hard-coding setwd in a script |
| List the objects in memory | ls() | |
| Remove an object | rm(x) | Restarting R is cleaner than rm(list = ls()) |
| Version of R and every loaded package | sessionInfo() | Paste this into any bug report or question |
| Time a piece of code | system.time(sum(1:1e7)) | |
| Make a script executable on Linux or macOS | #!/usr/bin/env Rscript | First line of the file, then chmod +x analysis.R |
| Comment | # one line | No block comment. Use # on each line |
Rscript is the command for scripts and scheduled jobs; R starts the interactive console. RStudio and Positron are editors that run the same R underneath, so everything on this page works in all of them.
Variables and types
| Task | Code | Notes |
|---|---|---|
| Assign a variable | name <- "Ada" | = works too at the top level. <- is the convention |
| Print a value | print(x) | At the console, typing the name on its own prints it |
| Print text without quotes or [1] | cat("Total:", 42, "\n") | Separates with a space and adds no newline unless you ask |
| Numbers | 42 3.14 1e6 | All doubles: typeof(42) is "double" |
| Whole number | 42L | The L makes an integer. 1:10 gives integers too |
| Text | "cat" 'cat' | The same. A character vector of length 1 |
| True and false | TRUE FALSE | T and F also work, but they are ordinary variables that can be overwritten |
| Missing value | NA | Unknown. length(NA) is 1 |
| No value at all | NULL | Nothing. length(NULL) is 0 |
| Not a number and infinity | 0 / 0 1 / 0 | NaN and Inf. is.na(NaN) is TRUE |
| What it is | class(x) | "numeric", "character", "data.frame", "factor" |
| How it is stored | typeof(x) | "double", "integer", "character", "list" |
| Test the type | is.numeric(x) is.character(x) | |
| Convert | as.numeric("42") as.character(42) | as.numeric("abc") is NA with a warning, not an error |
| To a whole number | as.integer(3.9) | 3. Truncates towards zero rather than rounding |
| Divide | 7 / 2 | 3.5 |
| Integer division | 7 %/% 2 | 3. Rounds down, so -7 %/% 2 is -4 |
| Remainder | 7 %% 3 | 1. Takes the sign of the divisor: -7 %% 3 is 2 |
| Power | 2^10 | 1024 |
| Round | round(3.14159, 2) | Halves go to the even number: round(2.5) is 2, round(3.5) is 4 |
| Significant figures | signif(123456, 2) | 120000 |
| Compare | x == 4 x != 4 x >= 10 | One TRUE or FALSE per element of x |
| And, or, not | x > 5 & x < 20 x < 5 | x > 20 !is.na(x) | Element by element. Use && and || inside if() |
| Is it missing | is.na(x) | x == NA is always NA, never TRUE |
| Default when NULL | user$nickname %||% "none" | R 4.4+. Replaces NULL only, not NA |
| Built-in constants | pi letters LETTERS month.name |
R has no scalars. 42 is a numeric vector of length 1, which is why the console prints [1] in front of it: the [1] is the position of the first value on that line.
Vectors
A vector holds values of one type in order. Most functions and operators work on the whole vector at once, so a loop is needed far less often than in other languages. The examples use x <- c(4, 8, 15, 16, 23, 42).
| Task | Code | Notes |
|---|---|---|
| Create a vector | x <- c(4, 8, 15, 16, 23, 42) | c() combines. Empty: c() is NULL, numeric(0) is an empty number vector |
| Whole numbers from a to b | 1:10 | 10:1 counts down |
| Sequence with a step | seq(0, 1, by = 0.25) | |
| Sequence of a given length | seq(0, 100, length.out = 5) | 0 25 50 75 100 |
| 1 to n, safely | seq_len(n) | 1:n gives 1 0 when n is 0. seq_len(0) is empty |
| Every position of a vector | seq_along(x) | Use it in for loops instead of 1:length(x) |
| Repeat | rep(c("a", "b"), times = 3) rep(c("a", "b"), each = 2) | a b a b a b, and a a b b |
| How many elements | length(x) | |
| Arithmetic on every element | x * 2 x + 100 | |
| Two vectors, element by element | c(1, 2, 3) + c(10, 20, 30) | 11 22 33 |
| Recycling | 1:6 + c(0, 100) | 1 102 3 104 5 106. The shorter vector repeats. Warns only when the lengths do not divide |
| Total, mean, median | sum(x) mean(x) median(x) | |
| Ignore missing values | mean(c(1, NA, 3), na.rm = TRUE) | 2. Without na.rm the answer is NA |
| Spread | sd(x) var(x) range(x) | range() returns the smallest and largest |
| Percentile | quantile(x, 0.9) | |
| Five-number summary and the mean | summary(x) | |
| Position of the largest | which.max(x) which.min(x) | |
| Running total | cumsum(x) | cumprod, cummax and cummin work the same way |
| Difference from the previous | diff(x) | One element shorter than x |
| Sort | sort(x) sort(x, decreasing = TRUE) | Drops NA unless na.last = TRUE |
| Positions that would sort it | order(x) | What you use to sort a data frame by a column |
| Reverse | rev(x) | |
| Unique values | unique(x) | duplicated(x) marks each repeat with TRUE |
| Count each value | table(c("a", "b", "a")) | |
| Is it in there | 15 %in% x | Works on a whole vector: c(15, 99) %in% x is TRUE FALSE |
| Is it not in there | 99 %notin% x | R 4.6+. !(99 %in% x) on older versions |
| Which positions match | which(x > 10) | 3 4 5 6 |
| Is any or every element true | any(x > 40) all(x > 0) | |
| Vector with names | ages <- c(ada = 36, grace = 45) | names(ages) reads or sets them |
| Set operations | union(1:3, 2:5) intersect(1:3, 2:5) setdiff(1:3, 2:5) | |
| Pick a value per element | ifelse(x > 10, "big", "small") | The vectorised if |
| Add to the end | x <- c(x, 99) | Copies the whole vector. In a long loop, create it at full length first |
| First and last few | head(x, 3) tail(x, 2) | |
| Mixing types | c(1, "a", TRUE) | Everything becomes text: "1" "a" "TRUE" |
Strings, factors and dates
| Task | Code | Notes |
|---|---|---|
| Join strings | paste("coding", "kitty") paste0("coding", "kitty") | "coding kitty" and "codingkitty" |
| Join a vector into one string | paste(words, collapse = ", ") | |
| Put values into a string | sprintf("%s scored %d", "Ada", 90L) | %s text, %d whole number, %f decimal |
| Two decimal places | sprintf("%.2f", 3.14159) | "3.14" |
| Thousands separator | format(1234567, big.mark = ",") | "1,234,567" |
| Pad a number with zeros | sprintf("%03d", 7L) | "007" |
| Length of a string | nchar("café") | 4. length("café") is 1: one string |
| Change case | toupper(s) tolower(s) | |
| Part of a string | substr(s, 1, 6) | Positions are inclusive and start at 1 |
| From a position to the end | substr(s, 8, NULL) | R 4.6+. substr(s, 8, nchar(s)) on older versions |
| Split | strsplit("a,b,c", ",") | Returns a list, one element per input string. strsplit(...)[[1]] for the first |
| Trim whitespace | trimws(" hi ") | |
| Does it match | grepl("an", words) | A regular expression. Add fixed = TRUE for plain text |
| Keep the matching strings | grepv("an", words) | R 4.5+. grep("an", words, value = TRUE) on older versions |
| Replace the first match, or all | sub("a", "o", "banana") gsub("a", "o", "banana") | "bonana" and "bonono" |
| Starts or ends with | startsWith(s, "cod") endsWith(s, ".csv") | |
| Categories with a fixed order | f <- factor(c("low", "high", "low"), levels = c("low", "high")) | Sorts and plots in level order, not alphabetically |
| The categories | levels(f) table(f) | |
| Factor of numbers back to numbers | as.numeric(as.character(factor(c("10", "5")))) | 10 5. as.numeric() on the factor alone gives the level codes, 1 2 |
| Text to a date | as.Date("2026-09-26") | ISO format by default. format = "%d/%m/%Y" for others |
| Format a date | format(as.Date("2026-09-26"), "%d %B %Y") | "26 September 2026" |
| Days between two dates | as.Date("2026-12-25") - as.Date("2026-09-26") | Time difference of 90 days. as.numeric() for the number |
| Today | Sys.Date() Sys.time() |
stringr, part of the tidyverse, wraps the same jobs with consistent names and the string always first: str_detect, str_replace_all, str_split, str_sub, str_pad. lubridate does the same for dates.
Indexing and subsetting
R counts from 1, not 0. Square brackets take positions, negative positions to leave out, names, or a TRUE/FALSE vector of the same length.
| Task | Code | Notes |
|---|---|---|
| First element | x[1] | x[0] is an empty vector, not an error |
| Last element | x[length(x)] tail(x, 1) | x[-1] is not the last element: it drops the first |
| Several positions | x[c(1, 3)] x[2:4] | |
| Everything except | x[-1] x[-c(1, 2)] | |
| Elements matching a condition | x[x > 10] | |
| Matching, with NAs dropped | x[which(x > 10)] | x[x > 10] keeps an NA for every NA in x |
| By name | ages["ada"] | |
| Past the end | x[10] | NA, not an error |
| Change one element | x[2] <- 0 | |
| Change every match | x[x > 20] <- 20 | |
| List element, as a list | user["name"] | One bracket keeps the container: a list of one |
| List element, the value | user[["name"]] user$name | |
| Column by a name held in a variable | col <- "score" df[[col]] | df$col looks for a column literally called col |
| One cell | df[2, "score"] | [row, column] |
| One row | df[2, ] | Leave the column side empty for all columns |
| One column, as a vector | df$score df[["score"]] df[, "score"] | |
| One column, as a data frame | df[, "score", drop = FALSE] | Tibbles never drop to a vector with [ |
| Several columns | df[, c("name", "score")] | |
| Rows matching a condition | df[df$score > 80, ] | A missing score gives a row of NAs. subset() and dplyr::filter() drop it |
| Rows and columns at once | df[df$team == "data", c("name", "score")] | |
| Matrix cell, row or column | m[2, 3] m[1, ] m[, 2] |
Lists
A list holds anything, each element its own type and length, lists included. Data frames, model results and parsed JSON are all lists underneath. The examples use user <- list(name = "Ada", age = 36, langs = c("R", "SQL")).
| Task | Code | Notes |
|---|---|---|
| Create a list | user <- list(name = "Ada", age = 36, langs = c("R", "SQL")) | Empty: list() |
| Read an element | user$name user[["name"]] | |
| Inside a nested element | user$langs[2] | "SQL" |
| Add or change an element | user$email <- "ada@example.com" | |
| Remove an element | user$email <- NULL | |
| Store NULL in an element | user["email"] <- list(NULL) | Assigning NULL with $ removes it instead |
| Element names | names(user) | |
| Is there an element called | "email" %in% names(user) | Safer than is.null(user$email): $ matches partial names |
| How many elements | length(user) | |
| See the structure | str(user) | The quickest way to understand an unfamiliar object |
| List of n empty slots | vector("list", 3) | Fill it in a loop with results[[i]] <- ... |
| Join two lists | c(list(a = 1), list(b = 2)) | |
| Override some settings | modifyList(list(a = 1, b = 2), list(b = 20)) | a = 1, b = 20. Works on nested lists too |
| Flatten to a vector | unlist(list(a = 1, b = 2:3)) | Named a, b1, b2 |
| Stack a list of data frames | do.call(rbind, frames) | dplyr::bind_rows(frames) also copes with columns that differ |
| Parse JSON into a list | jsonlite::fromJSON('{"name": "Ada", "langs": ["R", "SQL"]}') | jsonlite::toJSON() goes the other way |
Data frames
A data frame is a list of columns of the same length, one type per column. A tibble is the tidyverse's version: the same thing with stricter rules and tidier printing. The examples use a df with name, team and score columns.
| Task | Code | Notes |
|---|---|---|
| Create a data frame | df <- data.frame(name = c("Ada", "Grace", "Linus"), team = c("data", "data", "ops"), score = c(90, 85, NA)) | |
| Create a tibble | tibble::tibble(name = c("Ada", "Grace"), score = c(90, 85)) | |
| Rows and columns | nrow(df) ncol(df) dim(df) | |
| Column names | names(df) | |
| Overview | str(df) summary(df) | dplyr::glimpse(df) is the tidyverse version |
| First rows | head(df) head(df, 10) | |
| Type of every column | sapply(df, class) | |
| Add a column | df$passed <- df$score >= 88 | |
| Remove a column | df$passed <- NULL | |
| Rename a column | names(df)[names(df) == "score"] <- "points" | |
| Sort rows | df[order(df$score, decreasing = TRUE), ] | NA goes last. order(df$team, -df$score) sorts by two columns |
| Filter rows and pick columns | subset(df, score > 80, select = c(name, score)) | Drops rows where the condition is NA |
| Add rows | rbind(df, data.frame(name = "Alan", team = "ops", score = 70)) | Column names must match |
| Join two data frames | merge(df, teams, by = "team") | all.x = TRUE for a left join |
| Summary per group | aggregate(score ~ team, data = df, FUN = mean) | Drops rows with NA in the formula's columns first |
| Count per group | table(df$team) | |
| Missing values per column | colSums(is.na(df)) | |
| Drop rows with any NA | na.omit(df) df[complete.cases(df), ] | |
| Remove duplicate rows | unique(df) | |
| Built-in data to practise on | head(penguins) | R 4.5+. mtcars, iris and airquality work on any version |
Control flow
| Task | Code | Notes |
|---|---|---|
| If and else | if (n > 10) "big" else "small" | if is an expression, so it returns a value you can assign |
| If, else if, else | if (n > 10) { "big" } else if (n > 5) { "medium" } else { "small" } | At the top level of a script, else has to be on the same line as the closing } |
| If, for a whole vector | ifelse(x > 10, "big", "small") | if() needs exactly one TRUE or FALSE, and errors on a longer vector |
| Both conditions, stopping early | length(x) > 0 && x[1] > 0 | && and || need length-1 sides and error otherwise (R 4.3+) |
| Loop over positions | for (i in seq_along(x)) print(x[i]) | |
| Loop over values | for (w in words) cat(w, "\n") | |
| While | while (n > 0) n <- n - 1 | |
| Loop until break | repeat { n <- n + 1; if (n >= 10) break } | |
| Skip to the next pass, or stop | for (i in 1:5) { if (i == 2) next; if (i == 4) break; print(i) } | Prints 1 and 3 |
| Pick by value | switch(unit, cm = 0.01, m = 1, km = 1000, NA) | The last unnamed value is the default |
| Several values, one result | switch(day, sat = , sun = "weekend", "weekday") | An empty value falls through to the next |
| Raise an error | stop("score must be positive") | |
| Raise a warning | warning("using the default") | The code carries on |
| Check inputs in one line | stopifnot(is.numeric(x), length(x) > 0) | |
| Catch an error | tryCatch(stop("bad input"), error = function(e) conditionMessage(e)) | "bad input" |
| Catch a warning | tryCatch(log(-1), warning = function(w) NA) | log(-1) is NaN with a warning. This returns NA instead |
| Always run clean-up | tryCatch(risky(), finally = cat("done\n")) | Runs whether risky() succeeds or fails |
| Try and carry on | res <- try(log("a"), silent = TRUE) | inherits(res, "try-error") tells you whether it failed |
Functions and the pipe
| Task | Code | Notes |
|---|---|---|
| Define a function | add <- function(a, b = 1) a + b | b has a default. The last value is returned |
| Several lines | area <- function(w, h) { if (w < 0) return(NA); w * h } | return() is only needed to leave early |
| Call with names | add(b = 2, a = 1) | Named arguments can go in any order |
| Short anonymous function | \(x) x * 2 | R 4.1+. Short for function(x) x * 2 |
| Pass a function to another | sapply(1:3, \(i) i^2) | 1 4 9 |
| Any number of arguments | total <- function(...) sum(...) | list(...) captures them. Also passes arguments through to another function |
| Was an argument given | f <- function(a, b) if (missing(b)) a else a + b | |
| Only allow some values | f <- function(type = c("mean", "median")) match.arg(type) | f() gives "mean". f("mode") is an error listing the choices |
| Return several values | list(mean = mean(x), sd = sd(x)) | |
| Return without printing | invisible(x) | What functions called for their side effect usually return |
| Run clean-up on exit | on.exit(close(con), add = TRUE) | Inside a function. Runs even if the function errors |
| Call with a list of arguments | do.call(paste, list("a", "b", sep = "-")) | "a-b" |
| Pipe a value into a function | x |> sort() |> head(3) | R 4.1+. The same as head(sort(x), 3) |
| Pipe into another argument | penguins |> lm(body_mass ~ flipper_len, data = _) | R 4.2+. The _ must be a named argument |
| The tidyverse pipe | x %>% sort() %>% head(3) | magrittr's pipe, loaded with dplyr. |> needs no package and is the one to use now |
| See a function's code | add body(add) | Type the name without brackets |
Arguments are copied, in effect: changing x inside a function never changes the caller's x. Return the new value and assign it. <<- assigns outside the function, and is almost always a sign you should return something instead.
The apply family
These run a function on every element and collect the results, which replaces most loops. Pick one by what goes in and what you want back.
| Task | Code | Notes |
|---|---|---|
| Every element, results in a list | lapply(1:3, \(i) i^2) | Always a list, the same length as the input |
| Every element, results simplified | sapply(1:3, \(i) i^2) | A vector here, but can be a list or matrix. Fine at the console |
| Every element, result type guaranteed | vapply(words, nchar, integer(1)) | Errors if any result is not one integer. The safe choice inside functions |
| Two vectors in step | mapply(\(a, b) a * b, 1:3, 4:6) | 4 10 18. Map() does the same and always returns a list |
| Each row or column of a matrix | apply(m, 1, sum) apply(m, 2, sum) | 1 is rows, 2 is columns. rowSums and colMeans are faster |
| One value per group | tapply(penguins$body_mass, penguins$species, mean, na.rm = TRUE) | Extra arguments after the function are passed on to it |
| Split into groups, then apply | lapply(split(df$score, df$team), mean) | |
| Keep the elements that pass | Filter(\(v) v > 10, x) | 15 16 23 42 |
| First element that passes | Find(\(v) v > 10, x) | 15 |
| Combine step by step | Reduce(`+`, 1:4, accumulate = TRUE) | 1 3 6 10. Without accumulate, just the final 10 |
| Repeat an expression | replicate(3, mean(rnorm(10))) | Three means of fresh random numbers |
| purrr, the tidyverse version | purrr::map_dbl(1:3, \(i) i^2) | map() returns a list, map_dbl() and map_chr() name the type they return |
dplyr verbs
dplyr gives one verb per job. Every verb takes a data frame first and returns a new one, so they chain with the pipe. Run library(dplyr) first. The examples use penguins, which is built into R 4.5 and later.
| Task | Code | Notes |
|---|---|---|
| Keep rows | penguins |> filter(species == "Gentoo", body_mass > 5000) | A comma means and. Rows where the condition is NA are dropped |
| Keep rows matching either | penguins |> filter(species == "Adelie" | island == "Dream") | |
| Keep rows in a set | penguins |> filter(species %in% c("Adelie", "Gentoo")) | |
| Drop rows | penguins |> filter_out(island == "Torgersen") | dplyr 1.2+. Unlike filter(!...), keeps rows where the condition is NA |
| Pick columns | penguins |> select(species, island, body_mass) | |
| Drop a column | penguins |> select(-year) | |
| Pick columns by pattern or type | penguins |> select(starts_with("bill"), where(is.numeric)) | Also ends_with(), contains() and everything() |
| Rename | penguins |> rename(mass_g = body_mass) | New name on the left |
| Add or change a column | penguins |> mutate(mass_kg = body_mass / 1000) | Later columns in the same mutate() can use earlier ones |
| Column from a condition | penguins |> mutate(size = if_else(body_mass > 4500, "large", "small")) | NA stays NA. Stricter about types than ifelse() |
| Column from several conditions | penguins |> mutate(size = case_when(body_mass > 5000 ~ "large", body_mass > 3500 ~ "medium", .default = "small")) | First match wins. A missing body_mass gets the .default |
| Sort | penguins |> arrange(species, desc(body_mass)) | NA sorts last, even with desc() |
| Summarise | penguins |> summarise(avg = mean(body_mass, na.rm = TRUE), n = n()) | |
| Summarise per group | penguins |> summarise(avg = mean(body_mass, na.rm = TRUE), .by = species) | Nothing to ungroup afterwards |
| Group for several steps | penguins |> group_by(species) |> mutate(rank = min_rank(desc(body_mass))) |> ungroup() | group_by() stays attached until ungroup() |
| Count rows per group | penguins |> count(species, island, sort = TRUE) | |
| Unique combinations | penguins |> distinct(species, island) | |
| Top rows per group | penguins |> slice_max(body_mass, n = 1, by = species) | Ties are all kept. with_ties = FALSE for exactly n |
| First rows | penguins |> slice_head(n = 5) | |
| One column as a vector | penguins |> pull(body_mass) | |
| Same summary for several columns | penguins |> summarise(across(starts_with("bill"), \(v) mean(v, na.rm = TRUE))) | |
| Previous row's value | economics |> mutate(change = unemploy - lag(unemploy)) | lead() for the next row. economics comes with ggplot2 |
| Map old values to new | penguins |> mutate(sex = recode_values(sex, "female" ~ "F", "male" ~ "M")) | dplyr 1.2+. Unmatched values become NA. replace_values() keeps them |
| Left join | orders |> left_join(users, join_by(user_id == id)) | Every order, with the user's columns where there is a match |
| Inner and anti join | orders |> inner_join(users, join_by(user_id == id)) users |> anti_join(orders, join_by(id == user_id)) | Only matches, and users with no orders |
| Stack data frames | bind_rows(df, df) | Missing columns are filled with NA |
| Look at every column at once | glimpse(penguins) |
Coming from SQL? filter is WHERE, select is SELECT, arrange is ORDER BY, summarise with .by is GROUP BY, slice_head is LIMIT, and the joins keep their names. The SQL cheat sheet has the same jobs on the database side.
Reshaping with tidyr
tidyr changes the shape of a table without changing what is in it. Run library(tidyr) first. The examples use sales, a table with a region column and one column per quarter, q1 and q2.
| Task | Code | Notes |
|---|---|---|
| Wide to long | sales |> pivot_longer(c(q1, q2), names_to = "quarter", values_to = "amount") | One row per region and quarter. The shape ggplot2 wants |
| Long to wide | long |> pivot_wider(names_from = quarter, values_from = amount) | |
| Drop rows with a missing value | df |> drop_na(score) | With no columns named, drops a row with an NA anywhere |
| Replace missing values | df |> replace_na(list(score = 0)) | |
| Fill blanks from the row above | df |> fill(team) | |
| Split one column into several | people |> separate_wider_delim(name, delim = " ", names = c("first", "last")) | |
| Every combination, with the gaps | sales_long |> complete(region, quarter) | Missing combinations become rows of NA |
ggplot2 basics
ggplot2 builds a plot in layers: the data, aes() to map columns to x, y and colour, then a geom for each layer, joined with +. Run library(ggplot2) first.
| Task | Code | Notes |
|---|---|---|
| Scatter plot | ggplot(penguins, aes(flipper_len, body_mass)) + geom_point() | Rows with NA are left out, with a warning saying how many |
| Colour by a column | ggplot(penguins, aes(flipper_len, body_mass, colour = species)) + geom_point() | Adds the legend for you |
| Line chart | ggplot(economics, aes(date, unemploy)) + geom_line() | |
| Bar chart of counts | ggplot(penguins, aes(species)) + geom_bar() | Counts the rows for you |
| Bar chart of values | ggplot(totals, aes(species, avg)) + geom_col() | Bar heights you have already worked out |
| Histogram | ggplot(penguins, aes(body_mass)) + geom_histogram(binwidth = 250) | Always choose binwidth or bins. The default of 30 bins is a guess |
| Box plot | ggplot(penguins, aes(species, body_mass)) + geom_boxplot() | |
| Straight trend line | geom_smooth(method = "lm") | Add it after geom_point(). The band is the 95% confidence interval |
| One panel per group | facet_wrap(~ island) | facet_grid(sex ~ island) for a grid |
| Titles and axis labels | labs(title = "Flipper length vs mass", x = "Flipper (mm)", y = "Mass (g)", colour = "Species") | |
| Fixed colour for every point | geom_point(colour = "steelblue", alpha = 0.5) | Outside aes(). Inside aes() it would become a legend entry |
| Cleaner theme | theme_minimal() | Also theme_bw(), theme_classic() |
| Log scale | scale_y_log10() | |
| Zoom in | coord_cartesian(ylim = c(3000, 5000)) | ylim() removes the data outside, which changes box plots and trend lines |
| Horizontal bars | ggplot(penguins, aes(y = species)) + geom_bar() | Put the category on y |
| Save the last plot | ggsave("penguins.png", width = 7, height = 5, dpi = 300) | Width and height are in inches unless units = "cm" |
Reading and writing CSV
| Task | Code | Notes |
|---|---|---|
| Read a CSV | sales <- read.csv("sales.csv") | Base R. Text stays text since R 4.0 |
| Write a CSV | write.csv(sales, "out.csv", row.names = FALSE) | Without row.names = FALSE an unnamed column of row numbers is added |
| Read a CSV with readr | sales <- readr::read_csv("sales.csv") | Faster, returns a tibble and prints the column types it guessed |
| Set the column types | read_csv("sales.csv", col_types = cols(amount = col_double(), .default = col_character())) | Guessing is fine to explore. Pin the types in a script |
| Hide the column-type message | read_csv("sales.csv", show_col_types = FALSE) | |
| Values that mean missing | read_csv("sales.csv", na = c("", "NA", "-")) | |
| Skip lines at the top | read_csv("sales.csv", skip = 2) | |
| Semicolons and decimal commas | read_csv2("sales.csv") read.csv2("sales.csv") | The CSV Excel writes in much of Europe |
| Tabs or another separator | read_tsv("sales.tsv") read_delim("sales.txt", delim = "|") | |
| Read many files into one table | read_csv(files, id = "file") | files is a vector of paths. The id column says which file each row came from |
| Write a CSV with readr | write_csv(sales, "out.csv") | No row names. Missing values are written as NA; na = "" writes blanks |
| Save any R object exactly | saveRDS(sales, "sales.rds") sales <- readRDS("sales.rds") | Keeps types, factors and dates. Only R reads it |
| Read or write lines of text | readLines("notes.txt") writeLines(lines, "notes.txt") | |
| Build a path | file.path("data", "sales.csv") | "data/sales.csv" |
| Does the file exist | file.exists("sales.csv") | |
| Every CSV in a folder | list.files("data", pattern = "\\.csv$", full.names = TRUE) |
readxl reads Excel files with read_excel("sales.xlsx", sheet = 1), and is installed with the tidyverse. For JSON, jsonlite::read_json() and write_json().
Packages
| Task | Code | Notes |
|---|---|---|
| Install a package | install.packages("dplyr") | From CRAN. Once per machine, not in every script |
| Install several | install.packages(c("dplyr", "ggplot2", "readr", "tidyr")) | |
| Install the whole tidyverse | install.packages("tidyverse") | library(tidyverse) then loads the core packages in one go |
| Load a package | library(dplyr) | Once per session, at the top of the script |
| Use one function without loading | dplyr::n_distinct(c(1, 1, 2)) | Also says which one you mean when two packages share a name |
| Load only some functions | use("dplyr", c("filter", "select")) | R 4.5+. Nothing else from the package is attached |
| Is it installed | requireNamespace("dplyr", quietly = TRUE) | TRUE or FALSE, returned invisibly, so use it inside if (). Does not attach the package |
| Which version | packageVersion("dplyr") | |
| Update everything | update.packages(ask = FALSE) | |
| Uninstall | remove.packages("dplyr") | |
| Where packages are installed | .libPaths() | |
| A package's help index | help(package = "dplyr") | vignette(package = "dplyr") lists the long-form guides |
| Install from GitHub | pak::pak("tidyverse/dplyr") | Needs pak installed first |
| Start a project library | renv::init() | Gives the project its own package versions |
| Record the versions in use | renv::snapshot() | Writes renv.lock. Commit it |
| Install what renv.lock lists | renv::restore() | What someone cloning the project runs |
The install.packages, update.packages, pak and renv::restore rows were not run for this page: CRAN could not be reached from the build environment, so they are as documented rather than tested. Every other row in this group was run on R 4.6.1, renv::init and renv::snapshot with renv 1.1.4.
A dplyr pipeline, start to finish
Read a CSV, clean it, total it per group and write the result, which is most of what day-to-day analysis looks like. This runs as it is: it writes its own sample file first.
library(readr)
library(dplyr)
write_lines(c(
"date,region,product,units,price",
"2026-09-01,north,tea,12,2.50",
"2026-09-01,south,tea,8,2.50",
"2026-09-02,north,coffee,5,3.20",
"2026-09-02,south,coffee,,3.20",
"2026-09-03,north,tea,7,2.50",
"2026-09-03,south,cake,4,4.00"
), "sales.csv")
sales <- read_csv(
"sales.csv",
col_types = cols(
date = col_date(), units = col_integer(), price = col_double(),
.default = col_character()
)
)
report <- sales |>
filter(!is.na(units)) |>
mutate(revenue = units * price) |>
summarise(
orders = n(),
units = sum(units),
revenue = sum(revenue),
.by = c(region, product)
) |>
arrange(desc(revenue))
report
#> # A tibble: 4 × 5
#> region product orders units revenue
#> <chr> <chr> <int> <int> <dbl>
#> 1 north tea 2 19 47.5
#> 2 south tea 1 8 20
#> 3 north coffee 1 5 16
#> 4 south cake 1 4 16
write_csv(report, "report.csv")The empty units on 2 September is read as NA, and filter(!is.na(units))
drops that row before anything is summed. Without it, sum(units) for south
coffee would be NA, which spreads to every total it touches. Pinning
col_types means a stray letter in the price column gives a parsing-problems
warning and one NA, instead of readr quietly guessing that the whole column is
text.
library(dplyr) prints a note that it masks filter, lag, intersect and a few
others. That is expected: dplyr's versions now win when you call them by the plain
name, and stats::filter() is still there when you need it.
Base R and dplyr, side by side
The same jobs both ways, on penguins, and every pair was run to check it gives
the same result. Base R is always there with nothing to install; dplyr reads left
to right and chains with the pipe.
| Job | Base R | dplyr |
|---|---|---|
| Keep rows | subset(penguins, species == "Gentoo" & body_mass > 5000) | penguins |> filter(species == "Gentoo", body_mass > 5000) |
| Pick columns | penguins[, c("species", "body_mass")] | penguins |> select(species, body_mass) |
| Add a column | transform(penguins, mass_kg = body_mass / 1000) | penguins |> mutate(mass_kg = body_mass / 1000) |
| Rename a column | names(penguins)[names(penguins) == "body_mass"] <- "mass_g" | penguins |> rename(mass_g = body_mass) |
| Sort | penguins[order(penguins$species, -penguins$body_mass), ] | penguins |> arrange(species, desc(body_mass)) |
| Mean per group | aggregate(body_mass ~ species, data = penguins, FUN = mean) | penguins |> summarise(body_mass = mean(body_mass, na.rm = TRUE), .by = species) |
| Count per group | table(penguins$species) | penguins |> count(species) |
| Unique rows | unique(penguins[, c("species", "island")]) | penguins |> distinct(species, island) |
| Left join | merge(orders, users, by = "user_id", all.x = TRUE) | orders |> left_join(users, by = "user_id") |
| First rows | head(penguins, 5) | penguins |> slice_head(n = 5) |
aggregate with a formula drops rows with a missing value before it groups,
which is why it needs no na.rm. The dplyr version returns groups in the order
they first appear rather than sorted; add arrange(species) if the order matters.
Reshape, then plot
ggplot2 wants one row per point, so a table with a column per quarter is reshaped to long form first. This is the usual shape of a chart script.
library(tidyr)
library(ggplot2)
sales <- data.frame(
region = c("north", "south", "east"),
q1 = c(120, 95, 60),
q2 = c(135, 90, 80),
q3 = c(150, 110, 75)
)
long <- sales |>
pivot_longer(q1:q3, names_to = "quarter", values_to = "amount")
print(long, n = 4)
#> # A tibble: 9 × 3
#> region quarter amount
#> <chr> <chr> <dbl>
#> 1 north q1 120
#> 2 north q2 135
#> 3 north q3 150
#> 4 south q1 95
#> # ℹ 5 more rows
chart <- ggplot(long, aes(quarter, amount, colour = region, group = region)) +
geom_line(linewidth = 1) +
geom_point(size = 2) +
labs(title = "Sales by quarter", x = NULL, y = "Sales (£k)", colour = "Region") +
theme_minimal()
ggsave("sales.png", chart, width = 6, height = 4, dpi = 150)That saves a 900 by 600 pixel line chart with one coloured line per region.
group = region is the part people leave out. quarter is text, so ggplot2
splits the data by it as well as by colour, which leaves one point per group and
nothing to join: you get only points and the message
Each group consists of only one observation.
Passing chart to ggsave is safer than relying on the last plot, and inside a
loop or a function a plot is only drawn when you print() it.
Loop, apply or vectorise
Three ways to add 20% to every price. All three give the same answer; the last is the one to write.
prices <- c(tea = 2.5, coffee = 3.2, cake = 4)
# A for loop, with the result created at full length first
with_vat <- numeric(length(prices))
for (i in seq_along(prices)) {
with_vat[i] <- prices[i] * 1.2
}
names(with_vat) <- names(prices)
with_vat
#> tea coffee cake
#> 3.00 3.84 4.80
# The apply family: a function run on every element
vapply(prices, \(p) p * 1.2, numeric(1))
#> tea coffee cake
#> 3.00 3.84 4.80
# Vectorised: arithmetic already works on the whole vector
prices * 1.2
#> tea coffee cake
#> 3.00 3.84 4.80
# Where apply earns its place: a function that takes one value at a time
describe <- function(price) {
stopifnot(is.numeric(price), length(price) == 1)
if (price > 3) "premium" else "standard"
}
vapply(prices, describe, character(1))
#> tea coffee cake
#> "standard" "premium" "premium"describe uses if, which takes a single TRUE or FALSE, so it cannot be
handed the whole vector: describe(prices) stops at the stopifnot line. vapply
calls it once per price. For this particular job ifelse(prices > 3, "premium", "standard") is vectorised too, but most real functions are not that simple.
dplyr and SQL
The dplyr verbs map almost one to one onto SQL clauses, and the dbplyr package,
installed with the tidyverse, runs them against a database by writing the SQL for
you. show_query() prints what it wrote. This uses an in-memory SQLite database,
so it needs nothing else installed.
library(dplyr, warn.conflicts = FALSE)
library(DBI)
con <- dbConnect(RSQLite::SQLite(), ":memory:")
copy_to(con, penguins, "penguins")
query <- tbl(con, "penguins") |>
filter(!is.na(body_mass)) |>
summarise(avg_mass = mean(body_mass, na.rm = TRUE), n = n(), .by = species) |>
filter(n > 100) |>
arrange(desc(avg_mass))
show_query(query)
#> <SQL>
#> SELECT `species`, AVG(`body_mass`) AS `avg_mass`, COUNT(*) AS `n`
#> FROM `penguins`
#> WHERE (NOT((`body_mass` IS NULL)))
#> GROUP BY `species`
#> HAVING (COUNT(*) > 100.0)
#> ORDER BY `avg_mass` DESC
collect(query)
#> # A tibble: 2 × 3
#> species avg_mass n
#> <chr> <dbl> <int>
#> 1 Gentoo 5076. 123
#> 2 Adelie 3701. 151
dbDisconnect(con)A filter before summarise became WHERE, and the one after it became
HAVING, which is exactly the difference the
SQL cheat sheet explains.
Nothing is read into R until collect(). To write SQL yourself from R, DBI takes
? placeholders, as in
dbGetQuery(con, "SELECT * FROM users WHERE email = ?", params = list(email)), so
user input is never pasted into the query.
Coming from Python
The differences that trip people up in their first week, with the Python cheat sheet on the other side.
| Job | Python | R |
|---|---|---|
| First element | xs[0] | x[1] |
| Last element | xs[-1] | x[length(x)], because x[-1] drops the first |
| Assign | n = 5 | n <- 5 |
| True, false, nothing | True, False, None | TRUE, FALSE, NULL, and NA for a missing value |
| Define a function | def add(a, b=1): return a + b | add <- function(a, b = 1) a + b |
| Anonymous function | lambda x: x * 2 | \(x) x * 2 |
| Key-value pairs | {"name": "Ada"} | list(name = "Ada") |
| Loop | for x in xs: | for (x in xs) { } |
| Count to n | range(n), 0 to n - 1 | seq_len(n), 1 to n |
| Double every number | [v * 2 for v in xs] | x * 2 |
| Is it in there | x in xs | x %in% xs |
| Power, remainder, floor division | **, %, // | ^, %%, %/% |
| Put values in a string | f"{name} is {age}" | sprintf("%s is %d", name, age) |
| Import a library | import pandas as pd | library(dplyr) |
| Install a library | pip install pandas | install.packages("dplyr") |
Both languages round halves to the even number, so round(2.5) is 2 in each, and
both give -4 for -7 floor-divided by 2.
Gotchas
The mistakes almost everyone makes in their first month of R.
| Looks right | What actually happens | Do this instead |
|---|---|---|
x[-1] for the last element | Drops the first element | x[length(x)] or tail(x, 1) |
mean(x) on data with gaps | NA | mean(x, na.rm = TRUE) |
if (x == NA) | An error: x == NA is always NA | if (is.na(x)) |
if (x > 0) with a vector x | An error: the condition has length > 1 | any(x > 0), all(x > 0), or ifelse() for one result per element |
for (i in 1:length(x)) | Runs with i = 1 then 0 when x is empty | for (i in seq_along(x)) |
df[df$score > 80, ] | A row of NA for every missing score | subset(df, score > 80) or filter(df, score > 80) |
as.numeric(f) on a factor of numbers | The level codes, not the numbers | as.numeric(as.character(f)) |
user$na | Quietly returns user$name: $ matches partial names | user[["na"]], which only matches exactly |
sapply() inside a function | A vector most days, a list on the day one result is a different length | vapply() with the result type |
case_when(x > 5 ~ "high", .default = "low") | Missing x becomes "low" | Put is.na(x) ~ NA first |
select(species) after library(MASS) | unused argument: MASS's select masks dplyr's | Load MASS first, or write dplyr::select() |
geom_line() with text on the x axis | Only points, and Each group consists of only one observation | Add group = to aes() |
A ggplot made inside a for loop | Nothing is drawn | print() the plot |
write.csv(df, "out.csv") | An extra unnamed column of row numbers | row.names = FALSE, or readr::write_csv() |
0.1 + 0.2 == 0.3 | FALSE | isTRUE(all.equal(0.1 + 0.2, 0.3)) |
T and F for true and false | Work until something assigns T <- 0 | Spell out TRUE and FALSE |
sort() and order() pick an algorithm for you. For numbers, factors and logical
vectors the default is a radix sort,
which R's documentation says switches to an
insertion sort for fewer than
200 elements and to a counting sort
for integer vectors whose values span a range under 100,000. All three are on the site as
step-through visualisations.
Common questions
Which version of R does this cheat sheet cover?
R 4.6, the current version, checked against the 4.6.1 release, with dplyr 1.2.1, ggplot2 4.0.3, readr 2.2.0 and tidyr 1.3.2. Almost everything also works on R 4.1 and later. Anything newer says so in the notes column: the _ placeholder needs 4.2, the %||% operator 4.4, the built-in penguins data, grepv() and use() 4.5, and %notin% 4.6. filter_out() and recode_values() need dplyr 1.2. Run R --version, or R.version.string inside R, to see which version you have.
Should I learn base R or the tidyverse?
Both, in that order of need rather than of time. Vectors, indexing, lists and functions are base R and everything else is built on them, so the first sections of this page are unavoidable. For day-to-day data work, dplyr, tidyr, readr and ggplot2 are what most tutorials, courses and colleagues use, and they read more clearly than the base equivalents. The side-by-side table on this page shows the same jobs both ways, so you can read either style when you meet it.
What is the difference between <- and = in R?
For assigning a variable on its own line, none: x <- 5 and x = 5 do the same thing. Inside a function call they differ. median(x = 1:10) passes an argument called x and creates no variable, while median(x <- 1:10) would also assign x in your workspace. The R community and every major style guide use <- for assignment and = only for arguments, so that is what you will see in other people's code.
What is the difference between [ ], [[ ]] and $ in R?
Single brackets return the same kind of thing you started with: a list gives a smaller list and a data frame gives a smaller data frame. Double brackets reach inside and return one element itself, so user[["name"]] is the string, not a list holding it. $ is a shortcut for [[ ]] with a fixed name, as in df$score. Use [[ ]] when the name is in a variable, and be aware that $ on a list or data frame matches partial names, so user$na can quietly return user$name.
What is the difference between |> and %>%?
Both pass the value on the left into the function on the right, so x |> sort() and x %>% sort() both mean sort(x). |> is built into R from version 4.1 and needs no package. %>% comes from magrittr and is loaded with dplyr. They differ in the details: %>% uses a dot as its placeholder and allows it anywhere, while |> uses _ and only as a named argument, from R 4.2. New code can use |>, and older tutorials use %>%.
What is the difference between NA and NULL in R?
NA is a missing value that still takes up a place: c(1, NA, 3) has three elements, and length(NA) is 1. It means the value exists but is unknown. NULL is the absence of anything: c(1, NULL, 3) has two elements, and length(NULL) is 0. Test for them with is.na() and is.null(), never with ==. Data frames and vectors use NA for gaps; NULL is what a missing list element or an empty result gives back.
When should I use lapply, sapply or vapply?
lapply always returns a list, one element per input, so it is predictable and is the right choice when each result is something bigger than a single value. sapply tries to simplify that list into a vector or matrix, which is convenient at the console but can return a list when you expected a vector. vapply makes you state the type and length of each result, such as numeric(1), and errors if any result differs, so it is the safe choice inside functions. If you use purrr, map, map_dbl and map_chr follow the same idea.
Should I learn R or Python for data analysis?
Either will do the job, and many analysts use both. R was built for statistics, so models, statistical tests and publication-quality charts with ggplot2 need very little code, and it is common in research, health and academic statistics. Python is a general-purpose language, so the same code can also become a web service or an automation script, and it dominates machine learning. If your team or course already uses one, learn that one first. The Python cheat sheet on this site is grouped the same way as this page, so the two read side by side.
