Scripting and data cheat sheetR logo

R cheat sheet

R syntax on one page: vectors, data frames, indexing, the apply family, dplyr verbs, ggplot2 and reading CSV files, checked against R 4.6 and the tidyverse.

Last updated

R is a language built for statistics and working with data. Numbers come in vectors rather than one at a time, so a whole column is summed, filtered or plotted in a single line, and the tidyverse packages add one clear verb for each data job. The reference below is grouped by what you are trying to do, and the filter box searches all of it at once. Type dplyr to pull out the data-frame verbs, or 4.5 to see what arrived in recent releases.

Every snippet is checked against R 4.6, the current version, and was run on R 4.6.1 with dplyr 1.2.1, ggplot2 4.0.3, readr 2.2.0 and tidyr 1.3.2. Anything newer than R 4.1 says so in the notes column. Names like x, df and user are placeholders for your own, and the data examples use penguins, which ships with R from 4.5. Coming from spreadsheets or databases? The SQL cheat sheet covers the same filtering and grouping on the database side, and the Python cheat sheet is grouped the same way as this page.

Running R

TaskCommandNotes
Check which version you haveR --versionInside R: R.version.string
Run a scriptRscript analysis.R
Run one line without a fileRscript -e "print(2^10)"
Open the interactive consoleRLeave with q(). Answer n to saving the workspace
Run a script from inside Rsource("analysis.R")
Read the arguments passed to a scriptargs <- commandArgs(trailingOnly = TRUE)Rscript analysis.R sales.csv gives args[1] as "sales.csv"
Help for a function?mean help("mean")??regression searches every installed help page
Run the examples from a help pageexample(mean)
Where relative paths are read fromgetwd()setwd() changes it. Open the project folder instead of hard-coding setwd in a script
List the objects in memoryls()
Remove an objectrm(x)Restarting R is cleaner than rm(list = ls())
Version of R and every loaded packagesessionInfo()Paste this into any bug report or question
Time a piece of codesystem.time(sum(1:1e7))
Make a script executable on Linux or macOS#!/usr/bin/env RscriptFirst line of the file, then chmod +x analysis.R
Comment# one lineNo block comment. Use # on each line

Rscript is the command for scripts and scheduled jobs; R starts the interactive console. RStudio and Positron are editors that run the same R underneath, so everything on this page works in all of them.

Variables and types

TaskCodeNotes
Assign a variablename <- "Ada"= works too at the top level. <- is the convention
Print a valueprint(x)At the console, typing the name on its own prints it
Print text without quotes or [1]cat("Total:", 42, "\n")Separates with a space and adds no newline unless you ask
Numbers42 3.14 1e6All doubles: typeof(42) is "double"
Whole number42LThe L makes an integer. 1:10 gives integers too
Text"cat" 'cat'The same. A character vector of length 1
True and falseTRUE FALSET and F also work, but they are ordinary variables that can be overwritten
Missing valueNAUnknown. length(NA) is 1
No value at allNULLNothing. length(NULL) is 0
Not a number and infinity0 / 0 1 / 0NaN and Inf. is.na(NaN) is TRUE
What it isclass(x)"numeric", "character", "data.frame", "factor"
How it is storedtypeof(x)"double", "integer", "character", "list"
Test the typeis.numeric(x) is.character(x)
Convertas.numeric("42") as.character(42)as.numeric("abc") is NA with a warning, not an error
To a whole numberas.integer(3.9)3. Truncates towards zero rather than rounding
Divide7 / 23.5
Integer division7 %/% 23. Rounds down, so -7 %/% 2 is -4
Remainder7 %% 31. Takes the sign of the divisor: -7 %% 3 is 2
Power2^101024
Roundround(3.14159, 2)Halves go to the even number: round(2.5) is 2, round(3.5) is 4
Significant figuressignif(123456, 2)120000
Comparex == 4 x != 4 x >= 10One TRUE or FALSE per element of x
And, or, notx > 5 & x < 20 x < 5 | x > 20 !is.na(x)Element by element. Use && and || inside if()
Is it missingis.na(x)x == NA is always NA, never TRUE
Default when NULLuser$nickname %||% "none"R 4.4+. Replaces NULL only, not NA
Built-in constantspi letters LETTERS month.name

R has no scalars. 42 is a numeric vector of length 1, which is why the console prints [1] in front of it: the [1] is the position of the first value on that line.

Vectors

A vector holds values of one type in order. Most functions and operators work on the whole vector at once, so a loop is needed far less often than in other languages. The examples use x <- c(4, 8, 15, 16, 23, 42).

TaskCodeNotes
Create a vectorx <- c(4, 8, 15, 16, 23, 42)c() combines. Empty: c() is NULL, numeric(0) is an empty number vector
Whole numbers from a to b1:1010:1 counts down
Sequence with a stepseq(0, 1, by = 0.25)
Sequence of a given lengthseq(0, 100, length.out = 5)0 25 50 75 100
1 to n, safelyseq_len(n)1:n gives 1 0 when n is 0. seq_len(0) is empty
Every position of a vectorseq_along(x)Use it in for loops instead of 1:length(x)
Repeatrep(c("a", "b"), times = 3) rep(c("a", "b"), each = 2)a b a b a b, and a a b b
How many elementslength(x)
Arithmetic on every elementx * 2 x + 100
Two vectors, element by elementc(1, 2, 3) + c(10, 20, 30)11 22 33
Recycling1:6 + c(0, 100)1 102 3 104 5 106. The shorter vector repeats. Warns only when the lengths do not divide
Total, mean, mediansum(x) mean(x) median(x)
Ignore missing valuesmean(c(1, NA, 3), na.rm = TRUE)2. Without na.rm the answer is NA
Spreadsd(x) var(x) range(x)range() returns the smallest and largest
Percentilequantile(x, 0.9)
Five-number summary and the meansummary(x)
Position of the largestwhich.max(x) which.min(x)
Running totalcumsum(x)cumprod, cummax and cummin work the same way
Difference from the previousdiff(x)One element shorter than x
Sortsort(x) sort(x, decreasing = TRUE)Drops NA unless na.last = TRUE
Positions that would sort itorder(x)What you use to sort a data frame by a column
Reverserev(x)
Unique valuesunique(x)duplicated(x) marks each repeat with TRUE
Count each valuetable(c("a", "b", "a"))
Is it in there15 %in% xWorks on a whole vector: c(15, 99) %in% x is TRUE FALSE
Is it not in there99 %notin% xR 4.6+. !(99 %in% x) on older versions
Which positions matchwhich(x > 10)3 4 5 6
Is any or every element trueany(x > 40) all(x > 0)
Vector with namesages <- c(ada = 36, grace = 45)names(ages) reads or sets them
Set operationsunion(1:3, 2:5) intersect(1:3, 2:5) setdiff(1:3, 2:5)
Pick a value per elementifelse(x > 10, "big", "small")The vectorised if
Add to the endx <- c(x, 99)Copies the whole vector. In a long loop, create it at full length first
First and last fewhead(x, 3) tail(x, 2)
Mixing typesc(1, "a", TRUE)Everything becomes text: "1" "a" "TRUE"

Strings, factors and dates

TaskCodeNotes
Join stringspaste("coding", "kitty") paste0("coding", "kitty")"coding kitty" and "codingkitty"
Join a vector into one stringpaste(words, collapse = ", ")
Put values into a stringsprintf("%s scored %d", "Ada", 90L)%s text, %d whole number, %f decimal
Two decimal placessprintf("%.2f", 3.14159)"3.14"
Thousands separatorformat(1234567, big.mark = ",")"1,234,567"
Pad a number with zerossprintf("%03d", 7L)"007"
Length of a stringnchar("café")4. length("café") is 1: one string
Change casetoupper(s) tolower(s)
Part of a stringsubstr(s, 1, 6)Positions are inclusive and start at 1
From a position to the endsubstr(s, 8, NULL)R 4.6+. substr(s, 8, nchar(s)) on older versions
Splitstrsplit("a,b,c", ",")Returns a list, one element per input string. strsplit(...)[[1]] for the first
Trim whitespacetrimws(" hi ")
Does it matchgrepl("an", words)A regular expression. Add fixed = TRUE for plain text
Keep the matching stringsgrepv("an", words)R 4.5+. grep("an", words, value = TRUE) on older versions
Replace the first match, or allsub("a", "o", "banana") gsub("a", "o", "banana")"bonana" and "bonono"
Starts or ends withstartsWith(s, "cod") endsWith(s, ".csv")
Categories with a fixed orderf <- factor(c("low", "high", "low"), levels = c("low", "high"))Sorts and plots in level order, not alphabetically
The categorieslevels(f) table(f)
Factor of numbers back to numbersas.numeric(as.character(factor(c("10", "5"))))10 5. as.numeric() on the factor alone gives the level codes, 1 2
Text to a dateas.Date("2026-09-26")ISO format by default. format = "%d/%m/%Y" for others
Format a dateformat(as.Date("2026-09-26"), "%d %B %Y")"26 September 2026"
Days between two datesas.Date("2026-12-25") - as.Date("2026-09-26")Time difference of 90 days. as.numeric() for the number
TodaySys.Date() Sys.time()

stringr, part of the tidyverse, wraps the same jobs with consistent names and the string always first: str_detect, str_replace_all, str_split, str_sub, str_pad. lubridate does the same for dates.

Indexing and subsetting

R counts from 1, not 0. Square brackets take positions, negative positions to leave out, names, or a TRUE/FALSE vector of the same length.

TaskCodeNotes
First elementx[1]x[0] is an empty vector, not an error
Last elementx[length(x)] tail(x, 1)x[-1] is not the last element: it drops the first
Several positionsx[c(1, 3)] x[2:4]
Everything exceptx[-1] x[-c(1, 2)]
Elements matching a conditionx[x > 10]
Matching, with NAs droppedx[which(x > 10)]x[x > 10] keeps an NA for every NA in x
By nameages["ada"]
Past the endx[10]NA, not an error
Change one elementx[2] <- 0
Change every matchx[x > 20] <- 20
List element, as a listuser["name"]One bracket keeps the container: a list of one
List element, the valueuser[["name"]] user$name
Column by a name held in a variablecol <- "score" df[[col]]df$col looks for a column literally called col
One celldf[2, "score"][row, column]
One rowdf[2, ]Leave the column side empty for all columns
One column, as a vectordf$score df[["score"]] df[, "score"]
One column, as a data framedf[, "score", drop = FALSE]Tibbles never drop to a vector with [
Several columnsdf[, c("name", "score")]
Rows matching a conditiondf[df$score > 80, ]A missing score gives a row of NAs. subset() and dplyr::filter() drop it
Rows and columns at oncedf[df$team == "data", c("name", "score")]
Matrix cell, row or columnm[2, 3] m[1, ] m[, 2]

Lists

A list holds anything, each element its own type and length, lists included. Data frames, model results and parsed JSON are all lists underneath. The examples use user <- list(name = "Ada", age = 36, langs = c("R", "SQL")).

TaskCodeNotes
Create a listuser <- list(name = "Ada", age = 36, langs = c("R", "SQL"))Empty: list()
Read an elementuser$name user[["name"]]
Inside a nested elementuser$langs[2]"SQL"
Add or change an elementuser$email <- "ada@example.com"
Remove an elementuser$email <- NULL
Store NULL in an elementuser["email"] <- list(NULL)Assigning NULL with $ removes it instead
Element namesnames(user)
Is there an element called"email" %in% names(user)Safer than is.null(user$email): $ matches partial names
How many elementslength(user)
See the structurestr(user)The quickest way to understand an unfamiliar object
List of n empty slotsvector("list", 3)Fill it in a loop with results[[i]] <- ...
Join two listsc(list(a = 1), list(b = 2))
Override some settingsmodifyList(list(a = 1, b = 2), list(b = 20))a = 1, b = 20. Works on nested lists too
Flatten to a vectorunlist(list(a = 1, b = 2:3))Named a, b1, b2
Stack a list of data framesdo.call(rbind, frames)dplyr::bind_rows(frames) also copes with columns that differ
Parse JSON into a listjsonlite::fromJSON('{"name": "Ada", "langs": ["R", "SQL"]}')jsonlite::toJSON() goes the other way

Data frames

A data frame is a list of columns of the same length, one type per column. A tibble is the tidyverse's version: the same thing with stricter rules and tidier printing. The examples use a df with name, team and score columns.

TaskCodeNotes
Create a data framedf <- data.frame(name = c("Ada", "Grace", "Linus"), team = c("data", "data", "ops"), score = c(90, 85, NA))
Create a tibbletibble::tibble(name = c("Ada", "Grace"), score = c(90, 85))
Rows and columnsnrow(df) ncol(df) dim(df)
Column namesnames(df)
Overviewstr(df) summary(df)dplyr::glimpse(df) is the tidyverse version
First rowshead(df) head(df, 10)
Type of every columnsapply(df, class)
Add a columndf$passed <- df$score >= 88
Remove a columndf$passed <- NULL
Rename a columnnames(df)[names(df) == "score"] <- "points"
Sort rowsdf[order(df$score, decreasing = TRUE), ]NA goes last. order(df$team, -df$score) sorts by two columns
Filter rows and pick columnssubset(df, score > 80, select = c(name, score))Drops rows where the condition is NA
Add rowsrbind(df, data.frame(name = "Alan", team = "ops", score = 70))Column names must match
Join two data framesmerge(df, teams, by = "team")all.x = TRUE for a left join
Summary per groupaggregate(score ~ team, data = df, FUN = mean)Drops rows with NA in the formula's columns first
Count per grouptable(df$team)
Missing values per columncolSums(is.na(df))
Drop rows with any NAna.omit(df) df[complete.cases(df), ]
Remove duplicate rowsunique(df)
Built-in data to practise onhead(penguins)R 4.5+. mtcars, iris and airquality work on any version

Control flow

TaskCodeNotes
If and elseif (n > 10) "big" else "small"if is an expression, so it returns a value you can assign
If, else if, elseif (n > 10) { "big" } else if (n > 5) { "medium" } else { "small" }At the top level of a script, else has to be on the same line as the closing }
If, for a whole vectorifelse(x > 10, "big", "small")if() needs exactly one TRUE or FALSE, and errors on a longer vector
Both conditions, stopping earlylength(x) > 0 && x[1] > 0&& and || need length-1 sides and error otherwise (R 4.3+)
Loop over positionsfor (i in seq_along(x)) print(x[i])
Loop over valuesfor (w in words) cat(w, "\n")
Whilewhile (n > 0) n <- n - 1
Loop until breakrepeat { n <- n + 1; if (n >= 10) break }
Skip to the next pass, or stopfor (i in 1:5) { if (i == 2) next; if (i == 4) break; print(i) }Prints 1 and 3
Pick by valueswitch(unit, cm = 0.01, m = 1, km = 1000, NA)The last unnamed value is the default
Several values, one resultswitch(day, sat = , sun = "weekend", "weekday")An empty value falls through to the next
Raise an errorstop("score must be positive")
Raise a warningwarning("using the default")The code carries on
Check inputs in one linestopifnot(is.numeric(x), length(x) > 0)
Catch an errortryCatch(stop("bad input"), error = function(e) conditionMessage(e))"bad input"
Catch a warningtryCatch(log(-1), warning = function(w) NA)log(-1) is NaN with a warning. This returns NA instead
Always run clean-uptryCatch(risky(), finally = cat("done\n"))Runs whether risky() succeeds or fails
Try and carry onres <- try(log("a"), silent = TRUE)inherits(res, "try-error") tells you whether it failed

Functions and the pipe

TaskCodeNotes
Define a functionadd <- function(a, b = 1) a + bb has a default. The last value is returned
Several linesarea <- function(w, h) { if (w < 0) return(NA); w * h }return() is only needed to leave early
Call with namesadd(b = 2, a = 1)Named arguments can go in any order
Short anonymous function\(x) x * 2R 4.1+. Short for function(x) x * 2
Pass a function to anothersapply(1:3, \(i) i^2)1 4 9
Any number of argumentstotal <- function(...) sum(...)list(...) captures them. Also passes arguments through to another function
Was an argument givenf <- function(a, b) if (missing(b)) a else a + b
Only allow some valuesf <- function(type = c("mean", "median")) match.arg(type)f() gives "mean". f("mode") is an error listing the choices
Return several valueslist(mean = mean(x), sd = sd(x))
Return without printinginvisible(x)What functions called for their side effect usually return
Run clean-up on exiton.exit(close(con), add = TRUE)Inside a function. Runs even if the function errors
Call with a list of argumentsdo.call(paste, list("a", "b", sep = "-"))"a-b"
Pipe a value into a functionx |> sort() |> head(3)R 4.1+. The same as head(sort(x), 3)
Pipe into another argumentpenguins |> lm(body_mass ~ flipper_len, data = _)R 4.2+. The _ must be a named argument
The tidyverse pipex %>% sort() %>% head(3)magrittr's pipe, loaded with dplyr. |> needs no package and is the one to use now
See a function's codeadd body(add)Type the name without brackets

Arguments are copied, in effect: changing x inside a function never changes the caller's x. Return the new value and assign it. <<- assigns outside the function, and is almost always a sign you should return something instead.

The apply family

These run a function on every element and collect the results, which replaces most loops. Pick one by what goes in and what you want back.

TaskCodeNotes
Every element, results in a listlapply(1:3, \(i) i^2)Always a list, the same length as the input
Every element, results simplifiedsapply(1:3, \(i) i^2)A vector here, but can be a list or matrix. Fine at the console
Every element, result type guaranteedvapply(words, nchar, integer(1))Errors if any result is not one integer. The safe choice inside functions
Two vectors in stepmapply(\(a, b) a * b, 1:3, 4:6)4 10 18. Map() does the same and always returns a list
Each row or column of a matrixapply(m, 1, sum) apply(m, 2, sum)1 is rows, 2 is columns. rowSums and colMeans are faster
One value per grouptapply(penguins$body_mass, penguins$species, mean, na.rm = TRUE)Extra arguments after the function are passed on to it
Split into groups, then applylapply(split(df$score, df$team), mean)
Keep the elements that passFilter(\(v) v > 10, x)15 16 23 42
First element that passesFind(\(v) v > 10, x)15
Combine step by stepReduce(`+`, 1:4, accumulate = TRUE)1 3 6 10. Without accumulate, just the final 10
Repeat an expressionreplicate(3, mean(rnorm(10)))Three means of fresh random numbers
purrr, the tidyverse versionpurrr::map_dbl(1:3, \(i) i^2)map() returns a list, map_dbl() and map_chr() name the type they return

dplyr verbs

dplyr gives one verb per job. Every verb takes a data frame first and returns a new one, so they chain with the pipe. Run library(dplyr) first. The examples use penguins, which is built into R 4.5 and later.

TaskCodeNotes
Keep rowspenguins |> filter(species == "Gentoo", body_mass > 5000)A comma means and. Rows where the condition is NA are dropped
Keep rows matching eitherpenguins |> filter(species == "Adelie" | island == "Dream")
Keep rows in a setpenguins |> filter(species %in% c("Adelie", "Gentoo"))
Drop rowspenguins |> filter_out(island == "Torgersen")dplyr 1.2+. Unlike filter(!...), keeps rows where the condition is NA
Pick columnspenguins |> select(species, island, body_mass)
Drop a columnpenguins |> select(-year)
Pick columns by pattern or typepenguins |> select(starts_with("bill"), where(is.numeric))Also ends_with(), contains() and everything()
Renamepenguins |> rename(mass_g = body_mass)New name on the left
Add or change a columnpenguins |> mutate(mass_kg = body_mass / 1000)Later columns in the same mutate() can use earlier ones
Column from a conditionpenguins |> mutate(size = if_else(body_mass > 4500, "large", "small"))NA stays NA. Stricter about types than ifelse()
Column from several conditionspenguins |> mutate(size = case_when(body_mass > 5000 ~ "large", body_mass > 3500 ~ "medium", .default = "small"))First match wins. A missing body_mass gets the .default
Sortpenguins |> arrange(species, desc(body_mass))NA sorts last, even with desc()
Summarisepenguins |> summarise(avg = mean(body_mass, na.rm = TRUE), n = n())
Summarise per grouppenguins |> summarise(avg = mean(body_mass, na.rm = TRUE), .by = species)Nothing to ungroup afterwards
Group for several stepspenguins |> group_by(species) |> mutate(rank = min_rank(desc(body_mass))) |> ungroup()group_by() stays attached until ungroup()
Count rows per grouppenguins |> count(species, island, sort = TRUE)
Unique combinationspenguins |> distinct(species, island)
Top rows per grouppenguins |> slice_max(body_mass, n = 1, by = species)Ties are all kept. with_ties = FALSE for exactly n
First rowspenguins |> slice_head(n = 5)
One column as a vectorpenguins |> pull(body_mass)
Same summary for several columnspenguins |> summarise(across(starts_with("bill"), \(v) mean(v, na.rm = TRUE)))
Previous row's valueeconomics |> mutate(change = unemploy - lag(unemploy))lead() for the next row. economics comes with ggplot2
Map old values to newpenguins |> mutate(sex = recode_values(sex, "female" ~ "F", "male" ~ "M"))dplyr 1.2+. Unmatched values become NA. replace_values() keeps them
Left joinorders |> left_join(users, join_by(user_id == id))Every order, with the user's columns where there is a match
Inner and anti joinorders |> inner_join(users, join_by(user_id == id)) users |> anti_join(orders, join_by(id == user_id))Only matches, and users with no orders
Stack data framesbind_rows(df, df)Missing columns are filled with NA
Look at every column at onceglimpse(penguins)

Coming from SQL? filter is WHERE, select is SELECT, arrange is ORDER BY, summarise with .by is GROUP BY, slice_head is LIMIT, and the joins keep their names. The SQL cheat sheet has the same jobs on the database side.

Reshaping with tidyr

tidyr changes the shape of a table without changing what is in it. Run library(tidyr) first. The examples use sales, a table with a region column and one column per quarter, q1 and q2.

TaskCodeNotes
Wide to longsales |> pivot_longer(c(q1, q2), names_to = "quarter", values_to = "amount")One row per region and quarter. The shape ggplot2 wants
Long to widelong |> pivot_wider(names_from = quarter, values_from = amount)
Drop rows with a missing valuedf |> drop_na(score)With no columns named, drops a row with an NA anywhere
Replace missing valuesdf |> replace_na(list(score = 0))
Fill blanks from the row abovedf |> fill(team)
Split one column into severalpeople |> separate_wider_delim(name, delim = " ", names = c("first", "last"))
Every combination, with the gapssales_long |> complete(region, quarter)Missing combinations become rows of NA

ggplot2 basics

ggplot2 builds a plot in layers: the data, aes() to map columns to x, y and colour, then a geom for each layer, joined with +. Run library(ggplot2) first.

TaskCodeNotes
Scatter plotggplot(penguins, aes(flipper_len, body_mass)) + geom_point()Rows with NA are left out, with a warning saying how many
Colour by a columnggplot(penguins, aes(flipper_len, body_mass, colour = species)) + geom_point()Adds the legend for you
Line chartggplot(economics, aes(date, unemploy)) + geom_line()
Bar chart of countsggplot(penguins, aes(species)) + geom_bar()Counts the rows for you
Bar chart of valuesggplot(totals, aes(species, avg)) + geom_col()Bar heights you have already worked out
Histogramggplot(penguins, aes(body_mass)) + geom_histogram(binwidth = 250)Always choose binwidth or bins. The default of 30 bins is a guess
Box plotggplot(penguins, aes(species, body_mass)) + geom_boxplot()
Straight trend linegeom_smooth(method = "lm")Add it after geom_point(). The band is the 95% confidence interval
One panel per groupfacet_wrap(~ island)facet_grid(sex ~ island) for a grid
Titles and axis labelslabs(title = "Flipper length vs mass", x = "Flipper (mm)", y = "Mass (g)", colour = "Species")
Fixed colour for every pointgeom_point(colour = "steelblue", alpha = 0.5)Outside aes(). Inside aes() it would become a legend entry
Cleaner themetheme_minimal()Also theme_bw(), theme_classic()
Log scalescale_y_log10()
Zoom incoord_cartesian(ylim = c(3000, 5000))ylim() removes the data outside, which changes box plots and trend lines
Horizontal barsggplot(penguins, aes(y = species)) + geom_bar()Put the category on y
Save the last plotggsave("penguins.png", width = 7, height = 5, dpi = 300)Width and height are in inches unless units = "cm"

Reading and writing CSV

TaskCodeNotes
Read a CSVsales <- read.csv("sales.csv")Base R. Text stays text since R 4.0
Write a CSVwrite.csv(sales, "out.csv", row.names = FALSE)Without row.names = FALSE an unnamed column of row numbers is added
Read a CSV with readrsales <- readr::read_csv("sales.csv")Faster, returns a tibble and prints the column types it guessed
Set the column typesread_csv("sales.csv", col_types = cols(amount = col_double(), .default = col_character()))Guessing is fine to explore. Pin the types in a script
Hide the column-type messageread_csv("sales.csv", show_col_types = FALSE)
Values that mean missingread_csv("sales.csv", na = c("", "NA", "-"))
Skip lines at the topread_csv("sales.csv", skip = 2)
Semicolons and decimal commasread_csv2("sales.csv") read.csv2("sales.csv")The CSV Excel writes in much of Europe
Tabs or another separatorread_tsv("sales.tsv") read_delim("sales.txt", delim = "|")
Read many files into one tableread_csv(files, id = "file")files is a vector of paths. The id column says which file each row came from
Write a CSV with readrwrite_csv(sales, "out.csv")No row names. Missing values are written as NA; na = "" writes blanks
Save any R object exactlysaveRDS(sales, "sales.rds") sales <- readRDS("sales.rds")Keeps types, factors and dates. Only R reads it
Read or write lines of textreadLines("notes.txt") writeLines(lines, "notes.txt")
Build a pathfile.path("data", "sales.csv")"data/sales.csv"
Does the file existfile.exists("sales.csv")
Every CSV in a folderlist.files("data", pattern = "\\.csv$", full.names = TRUE)

readxl reads Excel files with read_excel("sales.xlsx", sheet = 1), and is installed with the tidyverse. For JSON, jsonlite::read_json() and write_json().

Packages

TaskCodeNotes
Install a packageinstall.packages("dplyr")From CRAN. Once per machine, not in every script
Install severalinstall.packages(c("dplyr", "ggplot2", "readr", "tidyr"))
Install the whole tidyverseinstall.packages("tidyverse")library(tidyverse) then loads the core packages in one go
Load a packagelibrary(dplyr)Once per session, at the top of the script
Use one function without loadingdplyr::n_distinct(c(1, 1, 2))Also says which one you mean when two packages share a name
Load only some functionsuse("dplyr", c("filter", "select"))R 4.5+. Nothing else from the package is attached
Is it installedrequireNamespace("dplyr", quietly = TRUE)TRUE or FALSE, returned invisibly, so use it inside if (). Does not attach the package
Which versionpackageVersion("dplyr")
Update everythingupdate.packages(ask = FALSE)
Uninstallremove.packages("dplyr")
Where packages are installed.libPaths()
A package's help indexhelp(package = "dplyr")vignette(package = "dplyr") lists the long-form guides
Install from GitHubpak::pak("tidyverse/dplyr")Needs pak installed first
Start a project libraryrenv::init()Gives the project its own package versions
Record the versions in userenv::snapshot()Writes renv.lock. Commit it
Install what renv.lock listsrenv::restore()What someone cloning the project runs

The install.packages, update.packages, pak and renv::restore rows were not run for this page: CRAN could not be reached from the build environment, so they are as documented rather than tested. Every other row in this group was run on R 4.6.1, renv::init and renv::snapshot with renv 1.1.4.

A dplyr pipeline, start to finish

Read a CSV, clean it, total it per group and write the result, which is most of what day-to-day analysis looks like. This runs as it is: it writes its own sample file first.

library(readr)
library(dplyr)
 
write_lines(c(
  "date,region,product,units,price",
  "2026-09-01,north,tea,12,2.50",
  "2026-09-01,south,tea,8,2.50",
  "2026-09-02,north,coffee,5,3.20",
  "2026-09-02,south,coffee,,3.20",
  "2026-09-03,north,tea,7,2.50",
  "2026-09-03,south,cake,4,4.00"
), "sales.csv")
 
sales <- read_csv(
  "sales.csv",
  col_types = cols(
    date = col_date(), units = col_integer(), price = col_double(),
    .default = col_character()
  )
)
 
report <- sales |>
  filter(!is.na(units)) |>
  mutate(revenue = units * price) |>
  summarise(
    orders = n(),
    units = sum(units),
    revenue = sum(revenue),
    .by = c(region, product)
  ) |>
  arrange(desc(revenue))
 
report
#> # A tibble: 4 × 5
#>   region product orders units revenue
#>   <chr>  <chr>    <int> <int>   <dbl>
#> 1 north  tea          2    19    47.5
#> 2 south  tea          1     8    20
#> 3 north  coffee       1     5    16
#> 4 south  cake         1     4    16
 
write_csv(report, "report.csv")

The empty units on 2 September is read as NA, and filter(!is.na(units)) drops that row before anything is summed. Without it, sum(units) for south coffee would be NA, which spreads to every total it touches. Pinning col_types means a stray letter in the price column gives a parsing-problems warning and one NA, instead of readr quietly guessing that the whole column is text.

library(dplyr) prints a note that it masks filter, lag, intersect and a few others. That is expected: dplyr's versions now win when you call them by the plain name, and stats::filter() is still there when you need it.

Base R and dplyr, side by side

The same jobs both ways, on penguins, and every pair was run to check it gives the same result. Base R is always there with nothing to install; dplyr reads left to right and chains with the pipe.

JobBase Rdplyr
Keep rowssubset(penguins, species == "Gentoo" & body_mass > 5000)penguins |> filter(species == "Gentoo", body_mass > 5000)
Pick columnspenguins[, c("species", "body_mass")]penguins |> select(species, body_mass)
Add a columntransform(penguins, mass_kg = body_mass / 1000)penguins |> mutate(mass_kg = body_mass / 1000)
Rename a columnnames(penguins)[names(penguins) == "body_mass"] <- "mass_g"penguins |> rename(mass_g = body_mass)
Sortpenguins[order(penguins$species, -penguins$body_mass), ]penguins |> arrange(species, desc(body_mass))
Mean per groupaggregate(body_mass ~ species, data = penguins, FUN = mean)penguins |> summarise(body_mass = mean(body_mass, na.rm = TRUE), .by = species)
Count per grouptable(penguins$species)penguins |> count(species)
Unique rowsunique(penguins[, c("species", "island")])penguins |> distinct(species, island)
Left joinmerge(orders, users, by = "user_id", all.x = TRUE)orders |> left_join(users, by = "user_id")
First rowshead(penguins, 5)penguins |> slice_head(n = 5)

aggregate with a formula drops rows with a missing value before it groups, which is why it needs no na.rm. The dplyr version returns groups in the order they first appear rather than sorted; add arrange(species) if the order matters.

Reshape, then plot

ggplot2 wants one row per point, so a table with a column per quarter is reshaped to long form first. This is the usual shape of a chart script.

library(tidyr)
library(ggplot2)
 
sales <- data.frame(
  region = c("north", "south", "east"),
  q1 = c(120, 95, 60),
  q2 = c(135, 90, 80),
  q3 = c(150, 110, 75)
)
 
long <- sales |>
  pivot_longer(q1:q3, names_to = "quarter", values_to = "amount")
 
print(long, n = 4)
#> # A tibble: 9 × 3
#>   region quarter amount
#>   <chr>  <chr>    <dbl>
#> 1 north  q1         120
#> 2 north  q2         135
#> 3 north  q3         150
#> 4 south  q1          95
#> # ℹ 5 more rows
 
chart <- ggplot(long, aes(quarter, amount, colour = region, group = region)) +
  geom_line(linewidth = 1) +
  geom_point(size = 2) +
  labs(title = "Sales by quarter", x = NULL, y = "Sales (£k)", colour = "Region") +
  theme_minimal()
 
ggsave("sales.png", chart, width = 6, height = 4, dpi = 150)

That saves a 900 by 600 pixel line chart with one coloured line per region. group = region is the part people leave out. quarter is text, so ggplot2 splits the data by it as well as by colour, which leaves one point per group and nothing to join: you get only points and the message Each group consists of only one observation. Passing chart to ggsave is safer than relying on the last plot, and inside a loop or a function a plot is only drawn when you print() it.

Loop, apply or vectorise

Three ways to add 20% to every price. All three give the same answer; the last is the one to write.

prices <- c(tea = 2.5, coffee = 3.2, cake = 4)
 
# A for loop, with the result created at full length first
with_vat <- numeric(length(prices))
for (i in seq_along(prices)) {
  with_vat[i] <- prices[i] * 1.2
}
names(with_vat) <- names(prices)
with_vat
#>    tea coffee   cake
#>   3.00   3.84   4.80
 
# The apply family: a function run on every element
vapply(prices, \(p) p * 1.2, numeric(1))
#>    tea coffee   cake
#>   3.00   3.84   4.80
 
# Vectorised: arithmetic already works on the whole vector
prices * 1.2
#>    tea coffee   cake
#>   3.00   3.84   4.80
 
# Where apply earns its place: a function that takes one value at a time
describe <- function(price) {
  stopifnot(is.numeric(price), length(price) == 1)
  if (price > 3) "premium" else "standard"
}
vapply(prices, describe, character(1))
#>        tea     coffee       cake
#> "standard"  "premium"  "premium"

describe uses if, which takes a single TRUE or FALSE, so it cannot be handed the whole vector: describe(prices) stops at the stopifnot line. vapply calls it once per price. For this particular job ifelse(prices > 3, "premium", "standard") is vectorised too, but most real functions are not that simple.

dplyr and SQL

The dplyr verbs map almost one to one onto SQL clauses, and the dbplyr package, installed with the tidyverse, runs them against a database by writing the SQL for you. show_query() prints what it wrote. This uses an in-memory SQLite database, so it needs nothing else installed.

library(dplyr, warn.conflicts = FALSE)
library(DBI)
 
con <- dbConnect(RSQLite::SQLite(), ":memory:")
copy_to(con, penguins, "penguins")
 
query <- tbl(con, "penguins") |>
  filter(!is.na(body_mass)) |>
  summarise(avg_mass = mean(body_mass, na.rm = TRUE), n = n(), .by = species) |>
  filter(n > 100) |>
  arrange(desc(avg_mass))
 
show_query(query)
#> <SQL>
#> SELECT `species`, AVG(`body_mass`) AS `avg_mass`, COUNT(*) AS `n`
#> FROM `penguins`
#> WHERE (NOT((`body_mass` IS NULL)))
#> GROUP BY `species`
#> HAVING (COUNT(*) > 100.0)
#> ORDER BY `avg_mass` DESC
 
collect(query)
#> # A tibble: 2 × 3
#>   species avg_mass     n
#>   <chr>      <dbl> <int>
#> 1 Gentoo     5076.   123
#> 2 Adelie     3701.   151
 
dbDisconnect(con)

A filter before summarise became WHERE, and the one after it became HAVING, which is exactly the difference the SQL cheat sheet explains. Nothing is read into R until collect(). To write SQL yourself from R, DBI takes ? placeholders, as in dbGetQuery(con, "SELECT * FROM users WHERE email = ?", params = list(email)), so user input is never pasted into the query.

Coming from Python

The differences that trip people up in their first week, with the Python cheat sheet on the other side.

JobPythonR
First elementxs[0]x[1]
Last elementxs[-1]x[length(x)], because x[-1] drops the first
Assignn = 5n <- 5
True, false, nothingTrue, False, NoneTRUE, FALSE, NULL, and NA for a missing value
Define a functiondef add(a, b=1): return a + badd <- function(a, b = 1) a + b
Anonymous functionlambda x: x * 2\(x) x * 2
Key-value pairs{"name": "Ada"}list(name = "Ada")
Loopfor x in xs:for (x in xs) { }
Count to nrange(n), 0 to n - 1seq_len(n), 1 to n
Double every number[v * 2 for v in xs]x * 2
Is it in therex in xsx %in% xs
Power, remainder, floor division**, %, //^, %%, %/%
Put values in a stringf"{name} is {age}"sprintf("%s is %d", name, age)
Import a libraryimport pandas as pdlibrary(dplyr)
Install a librarypip install pandasinstall.packages("dplyr")

Both languages round halves to the even number, so round(2.5) is 2 in each, and both give -4 for -7 floor-divided by 2.

Gotchas

The mistakes almost everyone makes in their first month of R.

Looks rightWhat actually happensDo this instead
x[-1] for the last elementDrops the first elementx[length(x)] or tail(x, 1)
mean(x) on data with gapsNAmean(x, na.rm = TRUE)
if (x == NA)An error: x == NA is always NAif (is.na(x))
if (x > 0) with a vector xAn error: the condition has length > 1any(x > 0), all(x > 0), or ifelse() for one result per element
for (i in 1:length(x))Runs with i = 1 then 0 when x is emptyfor (i in seq_along(x))
df[df$score > 80, ]A row of NA for every missing scoresubset(df, score > 80) or filter(df, score > 80)
as.numeric(f) on a factor of numbersThe level codes, not the numbersas.numeric(as.character(f))
user$naQuietly returns user$name: $ matches partial namesuser[["na"]], which only matches exactly
sapply() inside a functionA vector most days, a list on the day one result is a different lengthvapply() with the result type
case_when(x > 5 ~ "high", .default = "low")Missing x becomes "low"Put is.na(x) ~ NA first
select(species) after library(MASS)unused argument: MASS's select masks dplyr'sLoad MASS first, or write dplyr::select()
geom_line() with text on the x axisOnly points, and Each group consists of only one observationAdd group = to aes()
A ggplot made inside a for loopNothing is drawnprint() the plot
write.csv(df, "out.csv")An extra unnamed column of row numbersrow.names = FALSE, or readr::write_csv()
0.1 + 0.2 == 0.3FALSEisTRUE(all.equal(0.1 + 0.2, 0.3))
T and F for true and falseWork until something assigns T <- 0Spell out TRUE and FALSE

sort() and order() pick an algorithm for you. For numbers, factors and logical vectors the default is a radix sort, which R's documentation says switches to an insertion sort for fewer than 200 elements and to a counting sort for integer vectors whose values span a range under 100,000. All three are on the site as step-through visualisations.

Common questions

Which version of R does this cheat sheet cover?

R 4.6, the current version, checked against the 4.6.1 release, with dplyr 1.2.1, ggplot2 4.0.3, readr 2.2.0 and tidyr 1.3.2. Almost everything also works on R 4.1 and later. Anything newer says so in the notes column: the _ placeholder needs 4.2, the %||% operator 4.4, the built-in penguins data, grepv() and use() 4.5, and %notin% 4.6. filter_out() and recode_values() need dplyr 1.2. Run R --version, or R.version.string inside R, to see which version you have.

Should I learn base R or the tidyverse?

Both, in that order of need rather than of time. Vectors, indexing, lists and functions are base R and everything else is built on them, so the first sections of this page are unavoidable. For day-to-day data work, dplyr, tidyr, readr and ggplot2 are what most tutorials, courses and colleagues use, and they read more clearly than the base equivalents. The side-by-side table on this page shows the same jobs both ways, so you can read either style when you meet it.

What is the difference between <- and = in R?

For assigning a variable on its own line, none: x <- 5 and x = 5 do the same thing. Inside a function call they differ. median(x = 1:10) passes an argument called x and creates no variable, while median(x <- 1:10) would also assign x in your workspace. The R community and every major style guide use <- for assignment and = only for arguments, so that is what you will see in other people's code.

What is the difference between [ ], [[ ]] and $ in R?

Single brackets return the same kind of thing you started with: a list gives a smaller list and a data frame gives a smaller data frame. Double brackets reach inside and return one element itself, so user[["name"]] is the string, not a list holding it. $ is a shortcut for [[ ]] with a fixed name, as in df$score. Use [[ ]] when the name is in a variable, and be aware that $ on a list or data frame matches partial names, so user$na can quietly return user$name.

What is the difference between |> and %>%?

Both pass the value on the left into the function on the right, so x |> sort() and x %>% sort() both mean sort(x). |> is built into R from version 4.1 and needs no package. %>% comes from magrittr and is loaded with dplyr. They differ in the details: %>% uses a dot as its placeholder and allows it anywhere, while |> uses _ and only as a named argument, from R 4.2. New code can use |>, and older tutorials use %>%.

What is the difference between NA and NULL in R?

NA is a missing value that still takes up a place: c(1, NA, 3) has three elements, and length(NA) is 1. It means the value exists but is unknown. NULL is the absence of anything: c(1, NULL, 3) has two elements, and length(NULL) is 0. Test for them with is.na() and is.null(), never with ==. Data frames and vectors use NA for gaps; NULL is what a missing list element or an empty result gives back.

When should I use lapply, sapply or vapply?

lapply always returns a list, one element per input, so it is predictable and is the right choice when each result is something bigger than a single value. sapply tries to simplify that list into a vector or matrix, which is convenient at the console but can return a list when you expected a vector. vapply makes you state the type and length of each result, such as numeric(1), and errors if any result differs, so it is the safe choice inside functions. If you use purrr, map, map_dbl and map_chr follow the same idea.

Should I learn R or Python for data analysis?

Either will do the job, and many analysts use both. R was built for statistics, so models, statistical tests and publication-quality charts with ggplot2 need very little code, and it is common in research, health and academic statistics. Python is a general-purpose language, so the same code can also become a web service or an automation script, and it dominates machine learning. If your team or course already uses one, learn that one first. The Python cheat sheet on this site is grouped the same way as this page, so the two read side by side.

See all cheat sheets

Want this explained by a cat?

The videos cover the same ground in sixty seconds. If there is a tool you want a cheat sheet for next, ask.