Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

This upgraded R cheat sheet is organized around the work you actually do: set up a project, inspect and import data, clean and join it, visualize and model it, then make the result reproducible. It covers base R and tidyverse side by side, with practical checks for common silent errors. Version-sensitive notes were checked against sources available on August 18, 2026; the core examples work independently of a particular IDE.

Quick distinction: R is the language and computing environment; RStudio is an IDE that runs R; tidyverse is a collection of R packages; Quarto creates documents and other published content; and renv records project package dependencies.

The 60-second R map

R language and runtime
├── Base R: built-in objects, functions, statistics, graphics
├── Packages: additional tools, often installed from CRAN
├── IDE: RStudio or another environment for editing and running code
├── Quarto: reproducible reports, websites, presentations, and more
└── renv: project-specific package library and lockfile

RStudio is not R itself. You can run R without RStudio, and installing a new RStudio version does not by itself update the R installation it uses. The Posit RStudio user guide describes the IDE and its workflow.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Start with a project, not a pile of console commands

Install R first, then install an IDE separately if you want one. To check the active R installation and package versions from a session:

R.version.string
R.Version()
sessionInfo()
packageVersion("ggplot2")

R 4.6.1, “Happy Hop,” was listed as the latest R release on the official developer page, released June 24, 2026. R versions can change, so confirm the official R Developer Page when version currency matters. Posit listed RStudio 2026.07.1 in its release notes. These are separate version numbers for separate products.

Create an RStudio Project for a piece of work, keep scripts and data inside its folder structure, and prefer project-relative paths over machine-specific paths. For example, with the here package:

install.packages(c("tidyverse", "here", "renv"))
library(dplyr)

here::here("data", "raw", "sales.csv")

Use library() when you deliberately want to attach a package for interactive work. In reusable scripts, namespace calls such as dplyr::filter() make the source of a function explicit and help resolve name conflicts.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Core syntax and objects

x <- 10
y <- 20

x + y
x * y
x^2
x / y
x %% y       # remainder
x %/% y      # integer division

x == y
x != y
x >= y
TRUE & FALSE
TRUE | FALSE
!TRUE

<- is the idiomatic assignment operator in R. = is also valid in many assignment contexts and is used for named function arguments. Parentheses make complicated expressions easier to read and reduce surprises about operator precedence. Use TRUE and FALSE, not their reassignable abbreviations T and F.

# Comments begin with a hash sign
result <- mean(
  c(1, 2, 3),
  na.rm = TRUE
)

R’s everyday data structures differ in what they can hold and how they are indexed:

Structure Typical contents Access example
Atomic vector Values of one basic type x[1]
List Objects that may have different types x[[1]], x$name
Matrix or array Same-type values in two or more dimensions m[1, 2]
Data frame Rectangular columns, which may have different types df[["column"]]
Tibble A data-frame class designed for tidyverse workflows tbl$column, dplyr verbs
Factor Categorical values with defined levels levels(f)

Inspect before you transform:

class(x)
typeof(x)
length(x)
str(x)
attributes(x)
is.numeric(x)
is.character(x)
is.logical(x)
is.factor(x)
is.data.frame(x)

Missing and exceptional numeric values are not interchangeable. NA means a value is missing; NaN is a not-a-number result; NULL generally represents absence of an object or value; and Inf and -Inf are infinite numeric values. Test for missingness with is.na(x), not x == NA.

is.na(x)
anyNA(x)
mean(x, na.rm = TRUE)
na.omit(x)  # removes rows/elements with missing values

Removing missing values changes which observations contribute to a result. Check that choice against the question you are answering rather than applying na.rm = TRUE or na.omit() automatically.

Index and subset without surprises

x[1]
x[1:3]
x[-1]
x[x > 10]
x[c(TRUE, FALSE, TRUE)]

For data frames, [ selects rows and columns, [[ extracts one element or column, and $ is convenient when the column name is fixed in your code. [ usually preserves a container where possible, but a one-column result may simplify. Use drop = FALSE when you need to guarantee a data-frame-shaped result.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
df[1, 2]                         # row 1, column 2
df[1, ]                          # first row
df[, 2, drop = FALSE]            # keep a one-column data frame
df[["column"]]                   # extract a column
df$column                        # convenient fixed-name access
df[df$score > 80, , drop = FALSE]

Import, inspect, and export data

# Base R
df <- read.csv("data.csv")
df <- read.delim("data.tsv")
write.csv(df, "output.csv", row.names = FALSE)

# readr, commonly used with the tidyverse
df <- readr::read_csv("data.csv")
readr::write_csv(df, "output.csv")

# Native R serialization: preserves one R object's structure
saveRDS(df, "data.rds")
df <- readRDS("data.rds")

CSV is widely portable; RDS preserves an individual R object and its structure. RData can store multiple objects, but explicit one-object RDS files are often easier to reason about in scripts. For unfamiliar data, inspect types and values rather than assuming import guesses are correct.

str(df)
head(df)
tail(df)
summary(df)

# With dplyr
 dplyr::glimpse(df)
dplyr::count(df, group, sort = TRUE)
table(df$group, useNA = "ifany")

Remove the accidental leading space before dplyr::glimpse if copying the lines as a single snippet; R ignores leading whitespace in a command. A clean version is dplyr::glimpse(df). Never rely on objects left in the global environment by earlier console work: a reproducible script should create the objects it uses.

Clean and transform data

For a small task, base R can be direct. For a sequence of table operations, dplyr gives a readable vocabulary. Equivalent approaches are not always identical in edge cases, but both can be appropriate.

# Base R
df$age <- as.numeric(df$age)
adults <- subset(df, age >= 18)
adults$log_income <- log(adults$income)
aggregate(income ~ group, data = adults, FUN = mean, na.rm = TRUE)
# dplyr
clean <- df |>
  filter(age >= 18) |>
  mutate(log_income = log(income)) |>
  select(id, group, age, income, log_income) |>
  arrange(desc(income))
Task Common dplyr function
Keep rows matching a condition filter()
Keep or reorder columns select(), relocate()
Create or change columns mutate(), across()
Summarize values summarise() or summarize()
Sort rows arrange()
Rename, deduplicate, or count rename(), distinct(), count()
Group and ungroup group_by(), ungroup()
Choose rows by position slice()
Apply a conditional rule case_when(), if_else()
Use a fallback for missing values coalesce()
df |>
  group_by(group) |>
  summarise(
    n = n(),
    mean_income = mean(income, na.rm = TRUE),
    median_income = median(income, na.rm = TRUE),
    .groups = "drop"
  )
df |>
  mutate(status = case_when(
    score >= 90 ~ "Excellent",
    score >= 75 ~ "Good",
    TRUE ~ "Needs review"
  ))

To count missing values by column:

df |>
  summarise(across(everything(), ~ sum(is.na(.x))))

Join tables and validate the result

Joins match rows using keys. A left join keeps every row in the left table and adds matches from the right; an inner join keeps only matches; a full join keeps rows from both sides. The key risk is often not an error but an unexpected increase in rows when a key is duplicated.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
left_join(x, y, by = "id")
inner_join(x, y, by = "id")
right_join(x, y, by = "id")
full_join(x, y, by = "id")
semi_join(x, y, by = "id")  # keep x rows with a match in y
anti_join(x, y, by = "id")  # keep x rows without a match in y

# Explicit modern join specification
left_join(x, y, by = join_by(id))

Before and after a join, check key types, missing keys, duplicate keys, and row counts. Also check whitespace, capitalization, and formatting: "0012" and 12 are not necessarily the same key.

nrow(x)
nrow(y)

x |> count(id) |> filter(n > 1)
y |> count(id) |> filter(n > 1)

joined <- left_join(x, y, by = "id")
nrow(joined)

Duplicate keys can create multiple matched combinations and multiply rows. Do not impose a universal expected row count: decide what relationship the data should have, then check that the result matches it.

Reshape between wide and long data

Tidy data generally has one variable per column, one observation per row, and one value per cell. Pivot when the layout needed for analysis differs from the layout in the source.

long <- tidyr::pivot_longer(
  df,
  cols = starts_with("year_"),
  names_to = "year",
  values_to = "value"
)

wide <- tidyr::pivot_wider(
  long,
  names_from = year,
  values_from = value
)

Other useful tidyr tools include separate(), unite(), separate_wider_delim(), fill(), drop_na(), replace_na(), complete(), and unnest(). Check whether pivot keys are unique: widening data with multiple values per key may require an aggregation or a list-column.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Plot with ggplot2

ggplot2 builds a plot by combining data and aesthetic mappings with geometric layers and labels.

library(ggplot2)

ggplot(df, aes(x = age, y = income)) +
  geom_point() +
  labs(
    title = "Income by age",
    x = "Age",
    y = "Income"
  ) +
  theme_minimal()
Question or display Common layer
Relationship between numeric variables geom_point()
Trend or ordered observations geom_line()
Counts by category geom_bar()
Precomputed heights by category geom_col()
Distribution of a numeric variable geom_histogram(), geom_density()
Compare distributions across groups geom_boxplot(), geom_violin()
Add a fitted smooth geom_smooth()
Show values on a grid geom_tile()

geom_bar() counts observations by default. Use geom_col() when the data already contain the heights to draw. Put a constant style outside aes(); mappings inside aes() map a variable to a visual property.

ggplot(df, aes(x = age, y = income, color = group)) +
  geom_point() +
  facet_wrap(~ group) +
  scale_x_log10() +
  scale_y_continuous(labels = scales::comma) +
  labs(x = "Age", y = "Income", color = "Group")

Give units and legible labels, use palettes that remain interpretable for readers with color-vision differences and in grayscale, and remember that a polished plot is not evidence that the underlying statistical analysis is valid.

Summarize and model data

mean(x, na.rm = TRUE)
median(x, na.rm = TRUE)
sd(x, na.rm = TRUE)
var(x, na.rm = TRUE)
quantile(x, probs = c(.25, .5, .75), na.rm = TRUE)
cor(x, y, use = "complete.obs")

For a linear model, inspect more than the coefficient table:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
fit <- lm(y ~ x1 + x2, data = df)
summary(fit)
coef(fit)
confint(fit)
predict(fit, newdata = new_df)

par(mfrow = c(2, 2))
plot(fit)

A generalized linear model example for a binary outcome:

fit <- glm(
  outcome ~ age + treatment,
  data = df,
  family = binomial()
)
summary(fit)

These commands fit models; they do not establish that the study design, assumptions, missing-data handling, diagnostics, or interpretation are appropriate. Treat model checks as part of the analysis, not optional decoration.

Dates, strings, and factors

as.Date("2026-08-18")
format(Sys.Date(), "%Y-%m-%d")

lubridate::ymd("2026-08-18")
lubridate::year(date)
lubridate::month(date)
stringr::str_detect(x, "pattern")
stringr::str_replace(x, "old", "new")
stringr::str_extract(x, "\d+")
stringr::str_trim(x)
stringr::str_to_lower(x)
f <- factor(x)
levels(f)
forcats::fct_relevel(f, "Control", "Treatment")

Do not convert a factor of numeric-looking labels directly with as.numeric(f): that returns the factor’s internal level codes. Convert through character labels instead:

as.numeric(as.character(f))

Functions, iteration, and pipes

summarise_mean <- function(x, remove_missing = TRUE) {
  mean(x, na.rm = remove_missing)
}

add_tax <- function(price, rate = 0.2) {
  price * (1 + rate)
}

Use a for loop when its steps or state changes make the logic clearest. For repeated operations, base R offers lapply(), sapply(), and vapply(); purrr offers map(), typed variants such as map_dbl(), and walk() for side effects. Prefer a typed result when the output type matters, since sapply() may simplify results differently depending on the input.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
lapply(items, fun)
vapply(items, fun, numeric(1))
purrr::map(items, fun)
purrr::map_dbl(items, fun)
purrr::walk(items, fun)

purrr::map_dbl(list(1:3, 4:6), (x) mean(x))

R’s native pipe and the magrittr pipe express many of the same workflows but are not identical in every advanced use:

# Base R pipe
df |>
  filter(age >= 18) |>
  summarise(mean_age = mean(age))

# Magrittr pipe
df %>%
  filter(age >= 18) %>%
  summarise(mean_age = mean(age))

Use a named intermediate object when a pipeline is hard to debug, a result is reused, or a transformation marks a meaningful stage. A pipe is a readability tool, not a requirement.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Make analysis reproducible

Check the working directory and files when diagnosing a path problem:

getwd()
list.files()

setwd("path") can be useful interactively, but hard-coded working-directory changes often break shared scripts. Prefer an RStudio Project and project-relative paths. Record randomness and the session environment:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
set.seed(123)
sessionInfo()

Use renv to record and restore project package versions:

renv::init()
renv::snapshot()
renv::restore()
renv::status()

snapshot() records dependencies in a lockfile; restore() recreates the project package environment from that record. Commit the lockfile with the project. It does not automatically capture every system dependency, external file, database, or operating-system detail, so document those too. When upgrading R, package libraries may need to be installed or migrated for the new R version; Posit’s R upgrade guidance discusses upgrade considerations and multiple installations.

Quarto is a free option for reproducible reports, websites, presentations, and other publishing formats. An R code chunk in a .qmd file looks like this:

```{r}
summary(df)
```

Render from a terminal with quarto render report.qmd. See the Quarto site for supported formats and setup details. Current RStudio release notes describe Quarto workflow improvements, including PDF output choices; the exact options depend on the installed tools and release.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Debug common failures

Message or symptom Likely cause and next check
object 'x' not found It was not created, is misspelled, or is outside the current scope. Run the script from the beginning and inspect names.
could not find function The name may be wrong, or its package may not be installed or loaded. Check find("function_name") and namespace.
subscript out of bounds The requested position does not exist. Inspect dimensions, lengths, and index values.
non-numeric argument to binary operator An operand has the wrong type. Inspect with str() and convert only after checking its values.
replacement has ... rows The replacement length is incompatible with the target. Check dimensions and recycling assumptions.
Join has more rows than expected Look for duplicate keys, mismatched types, and many-to-many matches.
there is no package called ... Install it in the active R library, then verify the R version and library in use.

Useful inspection and recovery commands:

traceback()
warnings()
last.warning
debugonce(my_function)
browser()
recover()

sessionInfo()
find("function_name")
?function_name
example(function_name)

Package masking can make the same short function name refer to different functions. Check conflicts or qualify the call:

conflicts()
dplyr::filter(data, value > 0)
stats::filter(x)

Do not suppress warnings before understanding them. Inspect a small sample, structure, counts, and row totals at each major transformation.

Base R, tidyverse, or data.table?

Task Base R Tidyverse
Filter rows subset(), logical indexing filter()
Add or change a column df$new <- ... mutate()
Grouped summary aggregate() group_by() + summarise()
Join tables merge() *_join()
Reshape data reshape() pivot_longer(), pivot_wider()
Plot Base graphics ggplot()
Apply a function apply(), lapply() map(), across()

Choose base R for simple operations, minimal dependencies, and fundamentals; tidyverse when its consistent grammar and readable table pipelines suit the task or team; and data.table when its compact syntax, in-place updates, or performance characteristics fit the workload and your team understands its semantics. No one approach is universally best, and performance depends on the data, algorithm, memory allocation, I/O, and implementation. Tibbles print conservatively and fit tidyverse workflows; convert when needed with as.data.frame(tbl) or tibble::as_tibble(df).

What is upgraded in the current ecosystem?

Version-sensitive snapshot checked August 18, 2026. The major upgrade here is the reference’s coverage—from isolated syntax to a complete, reproducible workflow—not an official R release name. The R Developer Page listed R 4.6.1; Posit’s release notes listed RStudio 2026.07.1. The RStudio 2026.05 notes describe a faster Data Viewer with pinnable columns, a Summary sidebar, type-aware statistics, sparkline histograms, keyboard navigation, clipboard copying, and a default display limit increased from 50 to 200 columns. Those are IDE features, not changes to R’s language. Check the release notes for build-specific details and compatibility before upgrading.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Posit also publishes official visual references for base R, RStudio, and popular packages in its cheatsheet collection. Use them for quick lookup; this guide’s added value is connecting those commands into a workflow.

Printable quick reference

Need Start here
Inspect object str(x), class(x), summary(x)
Find missing values is.na(x), anyNA(x)
Import/export CSV readr::read_csv(), readr::write_csv()
Filter and transform filter(), mutate(), select()
Group and summarize group_by(), summarise()
Join and check left_join(), then inspect key duplicates and nrow()
Reshape pivot_longer(), pivot_wider()
Plot ggplot() + geom_* + labs()
Model lm() or glm(), then inspect diagnostics
Reproduce packages renv::snapshot(), renv::restore()
Debug str(), traceback(), sessionInfo()

Before running downloaded R scripts or loading serialized objects from unknown sources, inspect what you have and consider the trust of the source. An R script can execute code, and serialized objects should not be treated as automatically safe simply because their extension is .rds or .RData.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.