Free tools Windows power users keep installed
One-click scans. No signup required.
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
This upgraded R cheat sheet is organized around the work you actually do: set up a project, inspect and import data, clean and join it, visualize and model it, then make the result reproducible. It covers base R and tidyverse side by side, with practical checks for common silent errors. Version-sensitive notes were checked against sources available on August 18, 2026; the core examples work independently of a particular IDE.
Quick distinction: R is the language and computing environment; RStudio is an IDE that runs R; tidyverse is a collection of R packages; Quarto creates documents and other published content; and renv records project package dependencies.
The 60-second R map
R language and runtime
├── Base R: built-in objects, functions, statistics, graphics
├── Packages: additional tools, often installed from CRAN
├── IDE: RStudio or another environment for editing and running code
├── Quarto: reproducible reports, websites, presentations, and more
└── renv: project-specific package library and lockfile
RStudio is not R itself. You can run R without RStudio, and installing a new RStudio version does not by itself update the R installation it uses. The Posit RStudio user guide describes the IDE and its workflow.
Start with a project, not a pile of console commands
Install R first, then install an IDE separately if you want one. To check the active R installation and package versions from a session:
#1 Best Overall
R.version.string
R.Version()
sessionInfo()
packageVersion("ggplot2")
R 4.6.1, “Happy Hop,” was listed as the latest R release on the official developer page, released June 24, 2026. R versions can change, so confirm the official R Developer Page when version currency matters. Posit listed RStudio 2026.07.1 in its release notes. These are separate version numbers for separate products.
Create an RStudio Project for a piece of work, keep scripts and data inside its folder structure, and prefer project-relative paths over machine-specific paths. For example, with the here package:
install.packages(c("tidyverse", "here", "renv"))
library(dplyr)
here::here("data", "raw", "sales.csv")
Use library() when you deliberately want to attach a package for interactive work. In reusable scripts, namespace calls such as dplyr::filter() make the source of a function explicit and help resolve name conflicts.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchCore syntax and objects
x <- 10
y <- 20
x + y
x * y
x^2
x / y
x %% y # remainder
x %/% y # integer division
x == y
x != y
x >= y
TRUE & FALSE
TRUE | FALSE
!TRUE
<- is the idiomatic assignment operator in R. = is also valid in many assignment contexts and is used for named function arguments. Parentheses make complicated expressions easier to read and reduce surprises about operator precedence. Use TRUE and FALSE, not their reassignable abbreviations T and F.
# Comments begin with a hash sign
result <- mean(
c(1, 2, 3),
na.rm = TRUE
)
R’s everyday data structures differ in what they can hold and how they are indexed:
| Structure | Typical contents | Access example |
|---|---|---|
| Atomic vector | Values of one basic type | x[1] |
| List | Objects that may have different types | x[[1]], x$name |
| Matrix or array | Same-type values in two or more dimensions | m[1, 2] |
| Data frame | Rectangular columns, which may have different types | df[["column"]] |
| Tibble | A data-frame class designed for tidyverse workflows | tbl$column, dplyr verbs |
| Factor | Categorical values with defined levels | levels(f) |
Inspect before you transform:
class(x)
typeof(x)
length(x)
str(x)
attributes(x)
is.numeric(x)
is.character(x)
is.logical(x)
is.factor(x)
is.data.frame(x)
Missing and exceptional numeric values are not interchangeable. NA means a value is missing; NaN is a not-a-number result; NULL generally represents absence of an object or value; and Inf and -Inf are infinite numeric values. Test for missingness with is.na(x), not x == NA.
is.na(x)
anyNA(x)
mean(x, na.rm = TRUE)
na.omit(x) # removes rows/elements with missing values
Removing missing values changes which observations contribute to a result. Check that choice against the question you are answering rather than applying na.rm = TRUE or na.omit() automatically.
Index and subset without surprises
x[1]
x[1:3]
x[-1]
x[x > 10]
x[c(TRUE, FALSE, TRUE)]
For data frames, [ selects rows and columns, [[ extracts one element or column, and $ is convenient when the column name is fixed in your code. [ usually preserves a container where possible, but a one-column result may simplify. Use drop = FALSE when you need to guarantee a data-frame-shaped result.
df[1, 2] # row 1, column 2
df[1, ] # first row
df[, 2, drop = FALSE] # keep a one-column data frame
df[["column"]] # extract a column
df$column # convenient fixed-name access
df[df$score > 80, , drop = FALSE]
Import, inspect, and export data
# Base R
df <- read.csv("data.csv")
df <- read.delim("data.tsv")
write.csv(df, "output.csv", row.names = FALSE)
# readr, commonly used with the tidyverse
df <- readr::read_csv("data.csv")
readr::write_csv(df, "output.csv")
# Native R serialization: preserves one R object's structure
saveRDS(df, "data.rds")
df <- readRDS("data.rds")
CSV is widely portable; RDS preserves an individual R object and its structure. RData can store multiple objects, but explicit one-object RDS files are often easier to reason about in scripts. For unfamiliar data, inspect types and values rather than assuming import guesses are correct.
str(df)
head(df)
tail(df)
summary(df)
# With dplyr
dplyr::glimpse(df)
dplyr::count(df, group, sort = TRUE)
table(df$group, useNA = "ifany")
Remove the accidental leading space before dplyr::glimpse if copying the lines as a single snippet; R ignores leading whitespace in a command. A clean version is dplyr::glimpse(df). Never rely on objects left in the global environment by earlier console work: a reproducible script should create the objects it uses.
Clean and transform data
For a small task, base R can be direct. For a sequence of table operations, dplyr gives a readable vocabulary. Equivalent approaches are not always identical in edge cases, but both can be appropriate.
# Base R
df$age <- as.numeric(df$age)
adults <- subset(df, age >= 18)
adults$log_income <- log(adults$income)
aggregate(income ~ group, data = adults, FUN = mean, na.rm = TRUE)
# dplyr
clean <- df |>
filter(age >= 18) |>
mutate(log_income = log(income)) |>
select(id, group, age, income, log_income) |>
arrange(desc(income))
| Task | Common dplyr function |
|---|---|
| Keep rows matching a condition | filter() |
| Keep or reorder columns | select(), relocate() |
| Create or change columns | mutate(), across() |
| Summarize values | summarise() or summarize() |
| Sort rows | arrange() |
| Rename, deduplicate, or count | rename(), distinct(), count() |
| Group and ungroup | group_by(), ungroup() |
| Choose rows by position | slice() |
| Apply a conditional rule | case_when(), if_else() |
| Use a fallback for missing values | coalesce() |
df |>
group_by(group) |>
summarise(
n = n(),
mean_income = mean(income, na.rm = TRUE),
median_income = median(income, na.rm = TRUE),
.groups = "drop"
)
df |>
mutate(status = case_when(
score >= 90 ~ "Excellent",
score >= 75 ~ "Good",
TRUE ~ "Needs review"
))
To count missing values by column:
df |>
summarise(across(everything(), ~ sum(is.na(.x))))
Join tables and validate the result
Joins match rows using keys. A left join keeps every row in the left table and adds matches from the right; an inner join keeps only matches; a full join keeps rows from both sides. The key risk is often not an error but an unexpected increase in rows when a key is duplicated.
left_join(x, y, by = "id")
inner_join(x, y, by = "id")
right_join(x, y, by = "id")
full_join(x, y, by = "id")
semi_join(x, y, by = "id") # keep x rows with a match in y
anti_join(x, y, by = "id") # keep x rows without a match in y
# Explicit modern join specification
left_join(x, y, by = join_by(id))
Before and after a join, check key types, missing keys, duplicate keys, and row counts. Also check whitespace, capitalization, and formatting: "0012" and 12 are not necessarily the same key.
nrow(x)
nrow(y)
x |> count(id) |> filter(n > 1)
y |> count(id) |> filter(n > 1)
joined <- left_join(x, y, by = "id")
nrow(joined)
Duplicate keys can create multiple matched combinations and multiply rows. Do not impose a universal expected row count: decide what relationship the data should have, then check that the result matches it.
Reshape between wide and long data
Tidy data generally has one variable per column, one observation per row, and one value per cell. Pivot when the layout needed for analysis differs from the layout in the source.
Rank #3
long <- tidyr::pivot_longer(
df,
cols = starts_with("year_"),
names_to = "year",
values_to = "value"
)
wide <- tidyr::pivot_wider(
long,
names_from = year,
values_from = value
)
Other useful tidyr tools include separate(), unite(), separate_wider_delim(), fill(), drop_na(), replace_na(), complete(), and unnest(). Check whether pivot keys are unique: widening data with multiple values per key may require an aggregation or a list-column.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Plot with ggplot2
ggplot2 builds a plot by combining data and aesthetic mappings with geometric layers and labels.
library(ggplot2)
ggplot(df, aes(x = age, y = income)) +
geom_point() +
labs(
title = "Income by age",
x = "Age",
y = "Income"
) +
theme_minimal()
| Question or display | Common layer |
|---|---|
| Relationship between numeric variables | geom_point() |
| Trend or ordered observations | geom_line() |
| Counts by category | geom_bar() |
| Precomputed heights by category | geom_col() |
| Distribution of a numeric variable | geom_histogram(), geom_density() |
| Compare distributions across groups | geom_boxplot(), geom_violin() |
| Add a fitted smooth | geom_smooth() |
| Show values on a grid | geom_tile() |
geom_bar() counts observations by default. Use geom_col() when the data already contain the heights to draw. Put a constant style outside aes(); mappings inside aes() map a variable to a visual property.
ggplot(df, aes(x = age, y = income, color = group)) +
geom_point() +
facet_wrap(~ group) +
scale_x_log10() +
scale_y_continuous(labels = scales::comma) +
labs(x = "Age", y = "Income", color = "Group")
Give units and legible labels, use palettes that remain interpretable for readers with color-vision differences and in grayscale, and remember that a polished plot is not evidence that the underlying statistical analysis is valid.
Summarize and model data
mean(x, na.rm = TRUE)
median(x, na.rm = TRUE)
sd(x, na.rm = TRUE)
var(x, na.rm = TRUE)
quantile(x, probs = c(.25, .5, .75), na.rm = TRUE)
cor(x, y, use = "complete.obs")
For a linear model, inspect more than the coefficient table:
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →fit <- lm(y ~ x1 + x2, data = df)
summary(fit)
coef(fit)
confint(fit)
predict(fit, newdata = new_df)
par(mfrow = c(2, 2))
plot(fit)
A generalized linear model example for a binary outcome:
fit <- glm(
outcome ~ age + treatment,
data = df,
family = binomial()
)
summary(fit)
These commands fit models; they do not establish that the study design, assumptions, missing-data handling, diagnostics, or interpretation are appropriate. Treat model checks as part of the analysis, not optional decoration.
Rank #4
Dates, strings, and factors
as.Date("2026-08-18")
format(Sys.Date(), "%Y-%m-%d")
lubridate::ymd("2026-08-18")
lubridate::year(date)
lubridate::month(date)
stringr::str_detect(x, "pattern")
stringr::str_replace(x, "old", "new")
stringr::str_extract(x, "\d+")
stringr::str_trim(x)
stringr::str_to_lower(x)
f <- factor(x)
levels(f)
forcats::fct_relevel(f, "Control", "Treatment")
Do not convert a factor of numeric-looking labels directly with as.numeric(f): that returns the factor’s internal level codes. Convert through character labels instead:
as.numeric(as.character(f))
Functions, iteration, and pipes
summarise_mean <- function(x, remove_missing = TRUE) {
mean(x, na.rm = remove_missing)
}
add_tax <- function(price, rate = 0.2) {
price * (1 + rate)
}
Use a for loop when its steps or state changes make the logic clearest. For repeated operations, base R offers lapply(), sapply(), and vapply(); purrr offers map(), typed variants such as map_dbl(), and walk() for side effects. Prefer a typed result when the output type matters, since sapply() may simplify results differently depending on the input.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorslapply(items, fun)
vapply(items, fun, numeric(1))
purrr::map(items, fun)
purrr::map_dbl(items, fun)
purrr::walk(items, fun)
purrr::map_dbl(list(1:3, 4:6), (x) mean(x))
R’s native pipe and the magrittr pipe express many of the same workflows but are not identical in every advanced use:
# Base R pipe
df |>
filter(age >= 18) |>
summarise(mean_age = mean(age))
# Magrittr pipe
df %>%
filter(age >= 18) %>%
summarise(mean_age = mean(age))
Use a named intermediate object when a pipeline is hard to debug, a result is reused, or a transformation marks a meaningful stage. A pipe is a readability tool, not a requirement.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Make analysis reproducible
Check the working directory and files when diagnosing a path problem:
getwd()
list.files()
setwd("path") can be useful interactively, but hard-coded working-directory changes often break shared scripts. Prefer an RStudio Project and project-relative paths. Record randomness and the session environment:
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →set.seed(123)
sessionInfo()
Use renv to record and restore project package versions:
renv::init()
renv::snapshot()
renv::restore()
renv::status()
snapshot() records dependencies in a lockfile; restore() recreates the project package environment from that record. Commit the lockfile with the project. It does not automatically capture every system dependency, external file, database, or operating-system detail, so document those too. When upgrading R, package libraries may need to be installed or migrated for the new R version; Posit’s R upgrade guidance discusses upgrade considerations and multiple installations.
Quarto is a free option for reproducible reports, websites, presentations, and other publishing formats. An R code chunk in a .qmd file looks like this:
```{r}
summary(df)
```
Render from a terminal with quarto render report.qmd. See the Quarto site for supported formats and setup details. Current RStudio release notes describe Quarto workflow improvements, including PDF output choices; the exact options depend on the installed tools and release.
Recommended Free Tools
Debug common failures
| Message or symptom | Likely cause and next check |
|---|---|
object 'x' not found |
It was not created, is misspelled, or is outside the current scope. Run the script from the beginning and inspect names. |
could not find function |
The name may be wrong, or its package may not be installed or loaded. Check find("function_name") and namespace. |
subscript out of bounds |
The requested position does not exist. Inspect dimensions, lengths, and index values. |
non-numeric argument to binary operator |
An operand has the wrong type. Inspect with str() and convert only after checking its values. |
replacement has ... rows |
The replacement length is incompatible with the target. Check dimensions and recycling assumptions. |
| Join has more rows than expected | Look for duplicate keys, mismatched types, and many-to-many matches. |
there is no package called ... |
Install it in the active R library, then verify the R version and library in use. |
Useful inspection and recovery commands:
traceback()
warnings()
last.warning
debugonce(my_function)
browser()
recover()
sessionInfo()
find("function_name")
?function_name
example(function_name)
Package masking can make the same short function name refer to different functions. Check conflicts or qualify the call:
conflicts()
dplyr::filter(data, value > 0)
stats::filter(x)
Do not suppress warnings before understanding them. Inspect a small sample, structure, counts, and row totals at each major transformation.
Base R, tidyverse, or data.table?
| Task | Base R | Tidyverse |
|---|---|---|
| Filter rows | subset(), logical indexing |
filter() |
| Add or change a column | df$new <- ... |
mutate() |
| Grouped summary | aggregate() |
group_by() + summarise() |
| Join tables | merge() |
*_join() |
| Reshape data | reshape() |
pivot_longer(), pivot_wider() |
| Plot | Base graphics | ggplot() |
| Apply a function | apply(), lapply() |
map(), across() |
Choose base R for simple operations, minimal dependencies, and fundamentals; tidyverse when its consistent grammar and readable table pipelines suit the task or team; and data.table when its compact syntax, in-place updates, or performance characteristics fit the workload and your team understands its semantics. No one approach is universally best, and performance depends on the data, algorithm, memory allocation, I/O, and implementation. Tibbles print conservatively and fit tidyverse workflows; convert when needed with as.data.frame(tbl) or tibble::as_tibble(df).
What is upgraded in the current ecosystem?
Version-sensitive snapshot checked August 18, 2026. The major upgrade here is the reference’s coverage—from isolated syntax to a complete, reproducible workflow—not an official R release name. The R Developer Page listed R 4.6.1; Posit’s release notes listed RStudio 2026.07.1. The RStudio 2026.05 notes describe a faster Data Viewer with pinnable columns, a Summary sidebar, type-aware statistics, sparkline histograms, keyboard navigation, clipboard copying, and a default display limit increased from 50 to 200 columns. Those are IDE features, not changes to R’s language. Check the release notes for build-specific details and compatibility before upgrading.
Posit also publishes official visual references for base R, RStudio, and popular packages in its cheatsheet collection. Use them for quick lookup; this guide’s added value is connecting those commands into a workflow.
Printable quick reference
| Need | Start here |
|---|---|
| Inspect object | str(x), class(x), summary(x) |
| Find missing values | is.na(x), anyNA(x) |
| Import/export CSV | readr::read_csv(), readr::write_csv() |
| Filter and transform | filter(), mutate(), select() |
| Group and summarize | group_by(), summarise() |
| Join and check | left_join(), then inspect key duplicates and nrow() |
| Reshape | pivot_longer(), pivot_wider() |
| Plot | ggplot() + geom_* + labs() |
| Model | lm() or glm(), then inspect diagnostics |
| Reproduce packages | renv::snapshot(), renv::restore() |
| Debug | str(), traceback(), sessionInfo() |
Before running downloaded R scripts or loading serialized objects from unknown sources, inspect what you have and consider the trust of the source. An R script can execute code, and serialized objects should not be treated as automatically safe simply because their extension is .rds or .RData.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

