Skip to content

A primer on Rust

Published 25 July 2026

A running collection of Rust concepts worth knowing cold.

For a worked example of taking this somewhere else, see Calling Rust from Python with PyO3.

For lifetimes specifically — why impl<'a> Reader<'a> names 'a twice, and when a struct doesn’t need a lifetime parameter at all — see Lifetimes in Rust.

Cargo

Cargo is Rust’s build tool and package manager. It handles compiling your code, downloading and managing dependencies (crates), running tests, and publishing packages — all through one CLI.

Basic use cases

Create a new project

cargo new my_project      # new directory with a binary crate
cargo new my_lib --lib    # new directory with a library crate

Initialize Cargo in an existing directory

cargo init

Before creating anything, cargo init walks up the directory tree looking for an existing Cargo.toml — checking the current directory, then its parent, then its parent’s parent, and so on — to find out whether it’s being run inside an existing workspace. If it finds one, it treats the new package as a member of that workspace (and looks for the workspace’s src/bin layout) instead of creating a standalone package.

Build the project

cargo build            # debug build, output in target/debug
cargo build --release  # optimized build, output in target/release

Run the project

cargo run               # builds (if needed) and runs the binary
cargo run --release

Check for compile errors without producing a binary

cargo check

This is much faster than cargo build since it skips code generation — useful as a tight feedback loop while writing code.

Run tests

cargo test

Add a dependency

cargo add serde

This adds the crate to Cargo.toml and fetches it from crates.io.

Format and lint

cargo fmt     # auto-format code
cargo clippy  # lint for common mistakes and non-idiomatic patterns

Imagine you’re writing an essay. cargo fmt is like a teacher fixing your indentation, spacing, and punctuation so the essay looks neat. cargo clippy is like an editor saying, “This sentence is awkward,” “You repeated yourself,” or “There’s a simpler way to say this.”

Package names can’t start with a digit

cargo init picks the current directory name as the package name by default. If that name starts with a digit, Cargo refuses to create the package:

     Creating binary (application) package
error: invalid character `1` in package name: `1-basic-syntax`, the name cannot start with a digit
If you need a package name to not match the directory name, consider using --name flag.
If you need a binary with the name "1-basic-syntax", use a valid package name, and set the binary name to be different from the package. This can be done by setting the binary filename to `src/bin/1-basic-syntax.rs` or change the name in Cargo.toml with:

    [[bin]]
    name = "1-basic-syntax"
    path = "src/main.rs"

This comes up often if you’re organizing practice exercises into numbered folders (1-basic-syntax, 2-ownership, etc.) and run cargo init inside them directly. The fix is either of:

  • Pass an explicit package name: cargo init --name basic_syntax
  • Keep the directory name but override the binary name in Cargo.toml via a [[bin]] section, as shown above.

Hyphens in package names, underscores in code

Related trap, and one that’s usually explained badly. Say your package is named fizzbuzz-3:

[package]
name = "fizzbuzz-3"
use fizzbuzz-3::solve;   // syntax error
use fizzbuzz_3::solve;   // the only spelling that compiles

The reason is worth being precise about, because the usual “hyphens get converted to underscores” phrasing suggests a conversion you could opt out of:

  • The package name in Cargo.toml is just a string identifier for Cargo and crates.io. Hyphens are perfectly fine there.
  • Rust source identifiers can’t contain - at all. It isn’t a valid identifier character — the parser reads it as subtraction. fizzbuzz-3 in a use statement is a syntax error, not a valid-but-wrong name.
  • Cargo derives the crate’s Rust identifier from the package name by replacing every - with _. That derived identifier is what you reference in code, always and unconditionally.

So it isn’t “if I use a hyphen it gets converted.” You can never legally type - in that position in a .rs file. Cargo does the conversion on its side, package name → identifier, and your source has to already be written in the underscore form to compile at all.

The same asymmetry shows up with dependencies — hyphenated in the manifest, underscored at the use site:

[dependencies]
async-trait = "0.1"
tokio-util = "0.7"
use async_trait::async_trait;
use tokio_util::codec::Framed;

Hyphens are the more common convention for published crate names, so most of the ecosystem’s use statements are spelled differently from the crate names you search for on crates.io. That’s expected, not a mistake.

What is a workspace?

A Cargo workspace is a way to manage multiple related crates (packages) under a single umbrella.

  • A crate is one Rust package (library or executable).
  • A workspace is a collection of crates that are developed together.

For example, suppose you’re building a distributed database. Instead of one huge crate, you might organize it like this:

mydb/
├── Cargo.toml          <-- workspace
├── storage/
│   ├── Cargo.toml
│   └── src/
├── raft/
│   ├── Cargo.toml
│   └── src/
├── cli/
│   ├── Cargo.toml
│   └── src/
└── server/
    ├── Cargo.toml
    └── src/

Here storage is the storage engine, raft is the consensus library, cli is the command-line tool, and server is the database server — each its own crate.

The top-level Cargo.toml just declares the workspace:

[workspace]
members = [
    "storage",
    "raft",
    "cli",
    "server",
]

Notice there’s no [package] section here — the workspace root isn’t a crate itself.

Why use a workspace?

1. Shared target/ directory. Without a workspace, each crate compiles its dependencies separately (storage/target/, raft/target/, server/target/, …). With a workspace, everyone shares a single mydb/target/, so build artifacts aren’t duplicated and builds are much faster.

2. Shared dependencies. Suppose every crate uses Tokio. Without a workspace, each Cargo.toml repeats tokio = "1.48". With a workspace, the version is declared once:

# workspace Cargo.toml
[workspace.dependencies]
tokio = "1.48"

and each crate just writes:

tokio.workspace = true

Now every crate automatically uses the same version.

3. Build and test everything together. Instead of cd-ing into each crate to run cargo test, running cargo test from the workspace root tests every member crate.

4. Easy local dependencies. Suppose server uses raft. Inside server/Cargo.toml:

[dependencies]
raft = { path = "../raft" }

No publishing to crates.io needed.

src/bin vs. a workspace

These solve different problems.

src/bin gives you multiple executables inside one crate:

calculator/
├── Cargo.toml
└── src/
    ├── main.rs
    └── bin/
        ├── add.rs
        └── multiply.rs

There’s still one package. cargo run --bin add just picks one executable from it.

A workspace is different — add and multiply become completely separate crates:

calculator/
├── Cargo.toml      <-- workspace
├── add/
│   ├── Cargo.toml
│   └── src/
└── multiply/
    ├── Cargo.toml
    └── src/

Rule of thumb: one crate with multiple executables → use src/bin/. Multiple libraries or applications that should evolve independently but belong to the same project → use a workspace.

Many large projects — web frameworks, database engines, developer tools — are organized as workspaces, since they naturally split into several reusable crates.

Modules: Cargo compiles a graph, not a directory

I once spent an evening convinced cargo test was broken. I had written a dozen tests across src/numbers/basenum.rs, floatnum.rs and realnum.rs, and every run reported the same thing:

running 1 test
test tests::it_works ... ok

One test — the placeholder cargo new generates. Nothing I had written was being run, and it turned out nothing I had written was being compiled either.

The bad assumption was that Cargo walks src/ and picks up every .rs file it finds. It doesn’t. Compilation starts at the crate root — src/lib.rs for a library, src/main.rs for a binary — and from there the compiler follows mod declarations, and only mod declarations.

src/lib.rs
    │
    ▼
numbers/mod.rs
    │
    ├── basenum.rs
    ├── floatnum.rs
    └── realnum.rs

That tree only exists if somebody writes it down. lib.rs needs

pub mod numbers;

and numbers/mod.rs needs

pub mod basenum;
pub mod floatnum;
pub mod realnum;

Leave those three lines out and the files sit on disk as inert text. Their code isn’t compiled, their #[cfg(test)] blocks are never discovered, and cargo test faithfully reports the one test that was reachable. There’s no warning either, because from the compiler’s point of view there’s nothing to warn about — you never told it those files belong to the crate. A stray .rs file that no mod points at doesn’t even get parsed; you can fill it with nonsense and the build still passes.

Adding the missing declarations fixed it instantly. The model worth internalizing is that the directory layout only tells the compiler where to look for a module once you’ve declared it — it never decides whether the module exists. Rust’s module system is explicit all the way down, and that same explicitness is what governs compilation, visibility and test discovery.

Where a path starts: crate::, super::, self::

Declaring the modules is half of it. The other half is saying where a path begins, and there are three prefixes for that:

  • crate:: — start from the crate root (src/lib.rs for a library, src/main.rs for a binary)
  • super:: — start from the parent module, one level up from the current file
  • self:: — start from the current module

Take the tree above and say floatnum.rs needs a helper that lives next door in basenum.rs. These two lines reach the same function:

// in src/numbers/floatnum.rs
use crate::numbers::basenum::round_half_up;  // down from the root
use super::basenum::round_half_up;           // up one level, then across

self:: refers to the module you’re already in, so it’s almost always redundant — a bare basenum::round_half_up(..) inside numbers/mod.rs means exactly the same as self::basenum::round_half_up(..). It’s worth writing when the leading name would otherwise be ambiguous, most often because you’ve imported something with the same name from another crate.

// in src/numbers/mod.rs
pub mod basenum;

pub fn round(x: f64) -> f64 {
    self::basenum::round_half_up(x)   // unambiguously *my* basenum
}

A path with none of these prefixes is read as a crate name — which is why use std::io; works without ceremony and use numbers::basenum; doesn’t, unless there happens to be a dependency called numbers. (This is the 2018 edition rule; older code used a leading :: for the same job.)

Which to prefer is mostly about what you expect to change. crate:: paths read identically from every file in the crate and survive moving a file to a different depth; super:: is shorter and says “my sibling”, which is often what you actually mean, but it goes stale the moment the file moves.

/// vs. //!

Both are doc comments, and the ! tells you which direction they point. /// documents the item below it — a struct, a function, a module declaration — and desugars to the outer attribute #[doc = "..."]. //! documents the item containing it, so it desugars to the inner attribute #![doc = "..."] and only makes sense at the top of a module or of lib.rs, where it becomes the crate’s front page. That’s the whole difference: #[doc] attaches to what follows, #![doc] attaches to the thing you’re already inside.

//! Timestamp utilities.        // == #![doc = "Timestamp utilities."] on the crate

/// Returns the current timestamp.   // == #[doc = "..."] on `now`
pub fn now() -> u64 { /* ... */ }

Initializing objects

Rust has no constructors. There’s no special method the compiler calls, no new keyword, no initializer lists. A struct is built by writing out its fields, and everything else is convention layered on top of that.

struct Point {
    x: i32,
    y: i32,
}

let p = Point { x: 1, y: 2 };   // struct literal — this is the only primitive

Every field must be given a value. There’s no partial initialization and no implicit zeroing, which is why “uninitialized field” bugs don’t exist.

Field init shorthand. When a local variable has the same name as the field, write it once:

let x = 1;
let y = 2;
let p = Point { x, y };   // same as Point { x: x, y: y }

new() is a convention, not a keyword

String::new(), Vec::new(), HashMap::new() — these are just associated functions someone chose to name new. Nothing in the language treats them specially:

impl Point {
    fn new(x: i32, y: i32) -> Self {
        Point { x, y }
    }
}

let p = Point::new(1, 2);

Self is an alias for the type you’re in, so Point { x, y } and Self { x, y } are interchangeable here. Since it’s an ordinary function, you can have as many as you want with names that actually describe what they do — Point::origin(), Vec::with_capacity(n), String::from("hi").

The :: is doing the work described in the section on :: below: there’s no value to call a method on yet, so you go through the type.

Default and struct update syntax

For a struct with many fields where most have an obvious zero value, derive Default:

#[derive(Default)]
struct Config {
    host: String,      // ""
    port: u16,         // 0
    verbose: bool,     // false
    retries: u32,      // 0
}

let c = Config::default();

The derive requires every field’s type to be Default itself. To override defaults per-field, use #[derive(Default)] together with struct update syntax — ..expr fills in every field you didn’t name:

let c = Config {
    port: 8080,
    verbose: true,
    ..Default::default()
};

The .. must come last, and it moves out of the source value for any non-Copy field, so the thing on the right is usually a fresh Default::default() rather than a struct you still want to use.

If the derived zero values are wrong for your type, write the impl by hand:

impl Default for Config {
    fn default() -> Self {
        Config { host: "localhost".into(), port: 8080, verbose: false, retries: 3 }
    }
}

Other shapes

struct Meters(f64);          // tuple struct — constructed like a function call
let d = Meters(3.5);
let raw = d.0;

struct Marker;               // unit struct — the name is the value
let m = Marker;

enum Shape {                 // enum variants construct the same way
    Circle { r: f64 },
    Square(f64),
    Empty,
}
let s = Shape::Circle { r: 1.0 };

Tuple structs are the idiom for newtypes — wrapping a primitive to get a distinct type, so Meters and Seconds can’t be swapped by accident even though both are an f64 underneath.

Converting instead of constructing

The other common way to get a value is to convert an existing one. Implement From and you get Into for free:

impl From<(i32, i32)> for Point {
    fn from(t: (i32, i32)) -> Self {
        Point { x: t.0, y: t.1 }
    }
}

let p = Point::from((1, 2));
let p: Point = (1, 2).into();   // same thing, type driven by the annotation

This is why String::from("hi") and "hi".to_string() and "hi".into() all exist and all work — one From<&str> for String impl, reached three ways.

When the field list gets long

If construction has many optional parameters, the ecosystem’s answer is a builder: a separate struct that accumulates settings and produces the real value at the end.

let c = Config::builder().port(8080).verbose(true).build();

Each method takes self and returns Self, which is what makes the chaining work. Worth knowing it exists, but don’t reach for it early — ..Default::default() covers most of what a builder would, with none of the boilerplate.

Rule of thumb: struct literal by default, new() when there’s real work or invariants to enforce, Default + .. when most fields have sensible zeros, From when you’re converting, builder only when the parameter list is genuinely unwieldy.

if is an expression, not a statement

A natural first attempt at a “which is bigger” function, coming from C/C++/Java, looks like this:

fn bigger(a: i32, b: i32) -> bool {
    if (a > b) { True }
    False
}

This has a few problems, in increasing order of subtlety.

1. Booleans are lowercase. Rust uses true and false, not True and False.

2. Parentheses around the if condition are unnecessary. if a > b { ... } is preferred over if (a > b) { ... } — the compiler will even warn about it (unnecessary parentheses around if condition).

3. The real bug: an if without else has type (). Fixing the two issues above still doesn’t compile:

fn bigger(a: i32, b: i32) -> bool {
    if a > b {
        true
    }
    false
}

The error is a type mismatch: expected () , found bool. The reason is that if in Rust is an expression, and every expression has a type. An if with no else might not run its body at all, so the only type Rust can consistently give it is () — the unit type. Since the then branch here evaluates to true (a bool), it conflicts with the () the compiler expects from an else-less if.

This is different from if in C/C++/Java, where it’s a statement with no value at all, and if (a > b) { return true; } reads naturally. In Rust, if a > b { true } doesn’t mean “return true” — it means “the value of this branch is true,” and without an else, that value has nowhere consistent to go.

The rule: you cannot have a branch of an if (with no else) produce a value other than (), unless you use return to exit the function early instead of trying to yield a value from the if expression itself.

Three ways to fix it

// Fix 1: return early, so the if doesn't need to produce a value
fn bigger(a: i32, b: i32) -> bool {
    if a > b {
        return true;
    }
    false
}

// Fix 2: add an else, so the whole if expression has type bool
fn bigger(a: i32, b: i32) -> bool {
    if a > b {
        true
    } else {
        false
    }
}

// Fix 3: idiomatic — a > b is already a bool, no if needed
fn bigger(a: i32, b: i32) -> bool {
    a > b
}

Fix 3 is what most Rust code would actually use — since a > b already evaluates to a bool, wrapping it in an if is redundant. It also relies on another Rust rule: the last expression in a function is automatically returned if it doesn’t end with a semicolon.

Tests

#[cfg(test)]
mod tests {
    use super::*;

    #[test]
    fn it_biggers() {
        assert!(bigger(20, 10));
        assert!(!bigger(10, 20));
    }
}

Run with cargo test. The compiler error for the broken version looks like:

error[E0308]: mismatched types
  --> src/bin/02.rs:12:9
   |
11 | /     if a> b {
12 | |         true
   | |         ^^^^ expected `()`, found `bool`
13 | |     }
   | |_____- expected this to be `()`
   |
help: you might have meant to return this value
   |
12 |         return true;
   |         ++++++     +

The help line is the compiler nudging you toward Fix 1 — but Fix 3 is the one worth internalizing. Expression-oriented control flow, where if/match/blocks all produce values, is one of the biggest conceptual shifts coming from statement-oriented languages, and it’s worth getting comfortable with early.

Arrays, slices, and Vec

Most languages give you one list type. Rust gives you three, and the distinction between them is the first place ownership becomes concrete rather than theoretical.

Type Written Size Lives on Owns its data?
Array [T; N] Fixed at compile time Stack Yes
Vector Vec<T> Grows at runtime Heap Yes
Slice &[T] / &mut [T] Known at runtime Borrowed view No
let arr: [i32; 4] = [1, 2, 3, 4];   // length is part of the type
let vec: Vec<i32> = vec![1, 2, 3];  // can push/pop
let sl:  &[i32]   = &arr[1..3];     // a window into arr — [2, 3]

The length being part of the type is the key thing about arrays. [i32; 3] and [i32; 4] are different types, so a function taking [i32; 3] will not accept a four-element array. That’s why arrays are rare in function signatures and slices are everywhere.

The slice is the unifying view

A slice is a fat pointer: a pointer to some elements plus a length. It doesn’t care whether those elements came from an array, a Vec, or another slice — which makes &[T] the type you want in signatures:

fn sum(xs: &[i32]) -> i32 {
    xs.iter().sum()
}

let arr = [1, 2, 3];
let vec = vec![4, 5, 6];

sum(&arr);        // &[i32; 3] → &[i32]
sum(&vec);        // &Vec<i32> → &[i32]
sum(&vec[1..]);   // already a slice

All three calls work because of deref coercion (covered below). The practical rule: take &[T], return Vec<T>. Borrow the most general thing you can, hand back the thing you own.

Almost every method in this section is actually defined on [T] — the slice type — and Vec<T> and arrays inherit them by deref. That’s why vec.sort() and arr.sort() and slice.sort() are all the same function.

Creating them

let a = [0; 10];                       // ten zeros, [i32; 10]
let v = vec![0; 10];                   // ten zeros, Vec<i32>
let v = Vec::new();                    // empty, type inferred from later use
let v = Vec::with_capacity(1000);      // empty but pre-allocated
let v: Vec<i32> = (1..=5).collect();   // from an iterator
let v = "a b c".split(' ').collect::<Vec<_>>();

with_capacity matters more than it looks. A Vec that outgrows its buffer allocates a new one (typically double) and copies everything across. Amortised that’s O(1) per push, but if you know the final size, saying so up front avoids the copies entirely.

Reading elements

let v = vec![10, 20, 30];

v[0]              // 10 — panics if out of bounds
v.get(0)          // Some(&10)
v.get(99)         // None — the non-panicking version
v.first()         // Some(&10)
v.last()          // Some(&30)
v.len()           // 3
v.is_empty()      // false
v.contains(&20)   // true

Indexing with [] panics on an out-of-bounds access; .get() returns an Option. Use [] when an out-of-range index means your logic is broken, .get() when the index came from outside your control.

Minimum and maximum. Reach for the iterator methods rather than a hand-rolled loop:

let largest  = v.iter().max().unwrap();
let smallest = v.iter().min().unwrap();

They return Option<&T> — None for an empty slice — so you get a &i32 here. Dereference with * if you want an owned value.

That’s two traversals. If you want both in one pass, fold carries the pair along:

let (smallest, largest) = v.iter().fold(
    (i32::MAX, i32::MIN),
    |(min, max), &x| (min.min(x), max.max(x)),
);

The .min()/.max() inside the closure are i32’s own methods, not the iterator ones — worth knowing separately, since a.max(b) reads better than an explicit if.

Iterating. The three iterator constructors differ only in what they hand you:

for x in v.iter()     { }  // x: &i32     — borrow
for x in v.iter_mut() { *x *= 2; }  // x: &mut i32 — borrow mutably
for x in v            { }  // x: i32      — consumes v (IntoIterator)

for x in &v  { }  // shorthand for v.iter()
for &x in &v { }  // pattern destructures the reference, so x: i32

for (i, x) in v.iter().enumerate() { println!("{i}: {x}"); }

for &x in &v is the one that confuses people at first: the & on the left is a pattern that unwraps the reference, so x comes out as a plain i32 instead of &i32. It only works for Copy types.

Indexing by range — for i in 0..v.len() — works but is the least idiomatic option unless you genuinely need the index for something other than lookup.

The iterator methods worth becoming fluent in early: iter(), iter_mut(), into_iter(), enumerate(), map(), filter(), fold(), collect(), find(), any(), all(), max(), min(). These form the core of idiomatic Rust and show up throughout production code.

Growing and shrinking (Vec only)

These need ownership and a resizable buffer, so they exist on Vec but not on slices or arrays:

let mut v = vec![1, 2, 3];

v.push(4);              // [1, 2, 3, 4]
v.pop();                // Some(4), leaves [1, 2, 3]
v.insert(1, 99);        // [1, 99, 2, 3] — O(n), shifts everything right
v.remove(1);            // returns 99, [1, 2, 3] — O(n)
v.swap_remove(0);       // returns 1, O(1) but does not preserve order
v.extend([7, 8]);       // append an iterator's items
v.append(&mut other);   // move all of other's items in, emptying it
v.truncate(2);          // keep the first 2
v.clear();              // empty it, keeping the allocation
v.retain(|&x| x % 2 == 0);   // keep only elements matching a predicate
v.drain(1..3);          // remove a range, yielding the removed items

swap_remove is the one worth remembering: remove shifts every later element left, but swap_remove just moves the last element into the hole. If you don’t care about order, it turns an O(n) removal into O(1).

retain is the idiomatic filter-in-place. Reaching for v = v.into_iter().filter(...).collect() does the same thing with an extra allocation.

Sorting and searching

let mut v = vec![3, 1, 2];

v.sort();                              // ascending, stable
v.sort_unstable();                     // faster, no stability guarantee
v.sort_by(|a, b| b.cmp(a));            // descending
v.sort_by_key(|s| s.len());            // sort by a derived key
v.reverse();                           // in place

v.binary_search(&2);                   // Ok(idx) or Err(insert_position)
v.iter().position(|&x| x == 2);        // Some(idx) — linear scan
v.iter().find(|&&x| x > 1);            // Some(&value)

Two things trip people up here. First, sort needs Ord, which floats don’t implement — for f64 you need sort_by(|a, b| a.partial_cmp(b).unwrap()) or total_cmp. Second, binary_search returns a Result where the Err carries the index the element would go at, which is exactly what you want for an insertion:

let pos = v.binary_search(&x).unwrap_or_else(|e| e);
v.insert(pos, x);

Use sort_unstable by default for primitives — it’s faster and stability rarely matters. Use sort when equal elements have distinguishable identity you want preserved.

Deduplication

let mut v = vec![1, 1, 2, 2, 3, 1];
v.dedup();          // [1, 2, 3, 1] — only removes *consecutive* duplicates
v.sort();
v.dedup();          // [1, 2, 3] — sort first for true dedup

Slicing and splitting

let v = vec![1, 2, 3, 4, 5];

&v[1..3]            // [2, 3]
&v[..2]             // [1, 2]
&v[3..]             // [4, 5]

v.split_at(2);      // ([1, 2], [3, 4, 5])
v.chunks(2);        // [1,2], [3,4], [5] — non-overlapping
v.windows(2);       // [1,2], [2,3], [3,4], [4,5] — overlapping
v.split_first();    // Some((&1, &[2, 3, 4, 5]))
v.concat();         // flatten a Vec<Vec<T>> or Vec<&str>

windows is the one to reach for whenever you’re comparing adjacent pairs — checking whether a sequence is sorted, computing differences, and so on:

let sorted = v.windows(2).all(|w| w[0] <= w[1]);

Note that windows and chunks exist on slices, so they work on arrays and Vecs alike.

Mutating in place

let mut v = vec![1, 2, 3];

v.swap(0, 2);                    // [3, 2, 1]
v.fill(0);                       // [0, 0, 0]
v.rotate_left(1);                // shift elements left, wrapping
for x in v.iter_mut() { *x *= 2; }

Conversions

let v: Vec<i32> = arr.to_vec();          // array/slice → owned Vec
let s: &[i32]   = v.as_slice();          // Vec → slice (usually implicit)
let a: [i32; 3] = v.try_into().unwrap(); // Vec → array, fails if length differs
let joined = ["a", "b"].join("-");       // "a-b"

2D vectors

There’s no dedicated matrix type in the standard library. The usual idiom is a Vec of Vecs:

let mut grid = vec![vec![0; cols]; rows];
grid[r][c] = 1;

That’s rows separate heap allocations. For anything performance-sensitive, a flat Vec with manual indexing is meaningfully faster since it’s one allocation and one contiguous cache-friendly block:

let mut grid = vec![0; rows * cols];
grid[r * cols + c] = 1;

Quick reference

Task Method
Add to end / remove from end push() / pop()
Remove by index, keep order remove(i) — O(n)
Remove by index, order irrelevant swap_remove(i) — O(1)
Safe indexed access get(i) → Option<&T>
Smallest / largest element iter().min() / iter().max()
Both in one traversal fold()
Iterate with the index iter().enumerate()
Modify every element iter_mut()
Filter in place retain(|x| …)
Sort sort_unstable() / sort_by_key()
Remove duplicates sort() then dedup()
Search a sorted slice binary_search()
Search an unsorted slice iter().position()
Compare adjacent elements windows(2)
Fixed-size batches chunks(n)
Build from an iterator collect()

What is ::?

Coming from C++, :: looks like the scope resolution operator. Rust calls it the path separator, and it’s used to navigate into modules, types, traits, and enums. Take a line you’ll write constantly:

use std::io::{self, Write};

io::stdout().flush()?;

There are two different operators at work in that one expression.

:: goes through a namespace or a type. io::stdout() means “the function stdout inside the module io.” The full path is std::io::stdout(); the use std::io; import is what lets you shorten it.

. operates on a value you already have. io::stdout() returns a Stdout value, and .flush() is a method called on that value. So the expression switches modes halfway through:

io::stdout()   ::  module → function
    .flush()   .   value  → method

That’s the whole mental model: :: starts from a namespace or type, . operates on a particular value.

What can appear on the left of ::

Left side Example Meaning
Module std::io::stdout() Function inside a module
Type String::new() Associated function
Type u32::MAX Associated constant
Enum Option::Some(5) Enum variant
Trait <Dog as Animal>::sound() Associated item via a trait

The type cases are the ones that surprise people. In String::new(), String is not a module — it’s a type, and new is an associated function (Rust’s equivalent of a static method). It uses :: precisely because there’s no String value to call it on yet:

let mut s = String::new();  // :: — no value exists yet
s.push_str("hello");        // .  — operating on s

Generics slot into the path too. Vec::<i32>::new() has two ::: one to pin the generic parameter, one to reach the associated function.

Chaining the two operators is the pattern you’ll see everywhere — start from a type or module, then work on the value it produces:

String::new()   // :: create a String
    .trim()     // .  operate on it
    .len();     // .  operate on the result

Turbofish: ::<> vs. a type annotation

There are two places you can pin down a generic parameter, and I forget which one is idiomatic every single time. Take a struct:

struct Point<T> {
    x: T,
    y: T,
}

Both of these give you a Point<f64>:

let p: Point<f64> = Point { x: 3.0, y: 4.0 };   // annotation on the binding
let p = Point::<f64> { x: 3.0, y: 4.0 };        // turbofish on the constructor

The first annotates p and lets the compiler infer the generic on the right from that. The second — ::<...>, the turbofish — specifies the generic argument at the construction site, and p’s type follows from it.

For a let binding, prefer the annotation. Turbofish exists to solve a different problem: disambiguating a generic when there’s no left-hand side to infer from. Vec::<i32>::new() standing alone, or mid-chain calls like "42".parse::<i32>(), where the value never gets bound to a name you could annotate. If you already have a let, you already have a natural slot for the type, and the turbofish is redundant noise.

It also scales better once the generics get nested. Compare:

let cache: HashMap<String, Vec<Point<f64>>> = HashMap::new();   // reads like a signature
let cache = HashMap::<String, Vec<Point<f64>>>::new();          // ::< and >:: pile up

The annotation front-loads the type so you can scan it top-to-bottom; the turbofish version buries it inside a call expression between two pairs of colons.

Where turbofish really is the only option is a chain that ends in collect(), since collect is generic over its return type and there’s nothing else to infer from:

let nums = "1,2,3"
    .split(',')
    .map(str::parse::<i32>)
    .collect::<Result<Vec<_>, _>>()?;

Rule of thumb: binding to a variable, annotate the variable; stuck mid-expression, reach for the turbofish.

Turbofish picks a type argument. Its sibling — picking a trait implementation — is UFCS, below, and the two compose into <Type as Trait<T>>::method(...).

One aside on that Point<T>, since it bites once: both fields share the same parameter, so Point { x: 3.0, y: 4 } doesn’t compile — 4 infers as i32 while 3.0 infers as f64. Mixed types need two parameters, struct Point<T, U> { x: T, y: U }.

UFCS: which trait implementation?

Turbofish answers which generic type parameter. UFCS answers a different question — which implementation — and the two show up together often enough that it’s worth having both in the same chapter of your head.

UFCS stands for Universal Function Call Syntax, the name from RFC 132. It’s the name everyone still uses in conversation, but the Rust Reference calls the feature qualified paths, and the Book calls it fully qualified syntax. Same thing; if you go searching the official docs, “UFCS” won’t find much.

Method calls are sugar

Start from the thing you write every day. This:

dog.name();

is shorthand. The compiler resolves it to a real path, inserting however many &, &mut or * it needs to make a receiver fit:

Dog::name(&dog);

That second form is the one you can write by hand, and being able to write it by hand is the whole point — because sometimes the sugar is ambiguous and sometimes it picks the wrong thing.

Note that the desugared form takes the borrow explicitly. Method syntax auto-refs for you; path syntax does not.

Inherent methods win over trait methods

Give a type both an inherent method and a trait method with the same name:

trait Animal {
    fn name(&self);
}

struct Dog;

impl Animal for Dog {
    fn name(&self) {
        println!("Animal impl");
    }
}

impl Dog {
    fn name(&self) {
        println!("inherent impl");
    }
}

Then dog.name() prints inherent impl. This isn’t ambiguity and it isn’t an error — method resolution searches inherent impls before trait impls, so the inherent one shadows the trait one silently. That’s a deliberate rule (it lets a type “override” a trait method with a faster or more specific version), but it means the trait method becomes unreachable through dot syntax.

Three ways to name what you want:

let dog = Dog;

dog.name();                     // inherent impl — the shadowing winner
Dog::name(&dog);                // inherent impl — same resolution, spelled out
Animal::name(&dog);             // Animal impl — trait named, type inferred from the receiver
<Dog as Animal>::name(&dog);    // Animal impl — fully qualified, nothing left to infer

The last two are UFCS. Animal::name(&dog) is the short form and works whenever the receiver pins down the type. <Dog as Animal>::name(&dog) is the fully qualified form, and it’s the one you fall back to when the short form is still ambiguous.

Two traits, one method name

The case that forces your hand is two traits offering the same method to the same type:

trait First {
    fn foo(&self);
}

trait Second {
    fn foo(&self);
}

struct X;

impl First for X {
    fn foo(&self) { println!("First"); }
}

impl Second for X {
    fn foo(&self) { println!("Second"); }
}

Now x.foo() doesn’t compile:

error[E0034]: multiple applicable items in scope
  = note: candidate #1 is defined in an impl of the trait `First` for the type `X`
  = note: candidate #2 is defined in an impl of the trait `Second` for the type `X`
help: disambiguate the method call for candidate #1:
      First::foo(&x)

The compiler tells you the fix, and the fix is UFCS:

First::foo(&x);
Second::foo(&x);

// or, spelled out in full:
<X as First>::foo(&x);
<X as Second>::foo(&x);

Worth noting: this error only fires if both traits are actually in scope. Trait methods are only callable where the trait is imported, so a use that brings in a second trait can break a .foo() call that compiled fine yesterday — the classic “I added a dependency and an unrelated line stopped compiling” puzzle.

When there’s no receiver at all

The forms above still had a &x to infer the type from. Associated functions that take no self don’t, and that’s where the fully qualified form stops being optional:

trait Spawn {
    fn create() -> Self;
}

Spawn::create() is unresolvable — nothing in the expression says which implementor you want. You have to say it:

let d = <Dog as Spawn>::create();
let c = <Cat as Spawn>::create();

The same shape shows up for associated consts and associated types, which have no receiver either:

<i32 as Default>::default()
<u8 as num::Bounded>::max_value()
<Vec<i32> as IntoIterator>::Item        // in type position

You’ll write <T as Trait>::Assoc in generic code constantly, because inside a generic function T::Assoc is only unambiguous while exactly one bound provides it.

Turbofish inside UFCS

Now the two features meet. Make the trait itself generic, so a type can implement it more than once:

trait Convert<T> {
    fn convert(&self) -> T;
}

struct Foo;

impl Convert<i32> for Foo {
    fn convert(&self) -> i32 { 42 }
}

impl Convert<String> for Foo {
    fn convert(&self) -> String { "hello".to_string() }
}

Foo implements Convert twice — these are genuinely different impls, distinguished only by the trait’s generic parameter. foo.convert() is ambiguous, and so is Convert::convert(&foo), because naming the trait isn’t enough. You need the trait and its parameter:

let n = <Foo as Convert<i32>>::convert(&foo);       // 42
let s = <Foo as Convert<String>>::convert(&foo);    // "hello"

Reading it left to right:

<Foo as Convert<i32>>::convert(&foo)
 ^^^     ^^^^^^^ ^^^   ^^^^^^^
 type    trait   the    method
                 trait's
                 generic arg

Often you can shortcut by annotating the binding instead, exactly as with turbofish — let n: i32 = foo.convert(); resolves fine, because the expected type picks the impl. The fully qualified form is what you reach for when there’s no binding to annotate.

This is also the general shape of the standard library’s Into: x.into() is ambiguous the moment a type implements Into for more than one target, and the fix is either an annotation or <X as Into<Target>>::into(x).

Paths as values

One more place the path form matters: a method path with no call parentheses is a plain function value, so you can hand it to an iterator adapter.

let lens: Vec<usize> = words.iter().map(|s| s.len()).collect();
let lens: Vec<usize> = words.iter().map(|s| str::len(s)).collect();
let lens: Vec<usize> = words.iter().copied().map(str::len).collect();

That last line is the path form used directly as a function. str::parse::<i32> from the turbofish section is the same trick with a turbofish attached. The fully qualified version works too — <str as ToString>::to_string — and is occasionally the only way to name a specific trait’s method as a value.

One real gotcha it fixes

Auto-ref makes .clone() on a &&T do something surprising. If T is not Clone, then &T still is — references are always Copy — so the compiler happily resolves the call against the outer reference and hands you a &T back instead of the T you expected:

let outer: &&NotClone = &&value;
let c = outer.clone();          // c: &NotClone — a copied reference, no deep clone

No error, no warning, just a value that isn’t what you wanted. The fully qualified form refuses to auto-ref, so it tells you the truth:

let c = <NotClone as Clone>::clone(outer);   // error: NotClone doesn't implement Clone

Clippy has a lint for this (clone_double_ref), but the underlying lesson generalises: dot syntax is doing autoref and autoderef work on your behalf, and when a call resolves to something you didn’t expect, rewriting it as a qualified path is the fastest way to find out what the compiler actually picked.

The mental model

You write You’re answering
foo::<i32>() which generic type argument?
Trait::foo(&x) which trait?
<Type as Trait>::foo(&x) which trait, on which type?
<Type as Trait<T>>::foo(&x) which trait, on which type, at which generic argument?

Reach for the shortest form that compiles. The fully qualified one is verbose on purpose — it’s the escape hatch, not the default — but knowing it exists turns a class of “multiple applicable items in scope” errors from a wall into a one-line fix.

PhantomData: a type parameter that carries no data

Sometimes you want a generic parameter purely as a tag — something the type system can distinguish on, with no runtime value behind it. Rust won’t let you declare struct Distance<T> { value: f64 } directly; every type parameter has to be used by some field, or the compiler rejects it as unconstrained. std::marker::PhantomData<T> is the zero-sized placeholder that satisfies that requirement.

The unit-of-measure example is the one that makes it click:

use std::marker::PhantomData;

struct Meters;
struct Feet;

struct Distance<T> {
    value: f64,
    unit: PhantomData<T>,
}

fn main() {
    let dist_m: Distance<Meters> = Distance {
        value: 10.0,
        unit: PhantomData,
    };

    let dist_f: Distance<Feet> = Distance {
        value: 3.0,
        unit: PhantomData,
    };

    println!("Meters: {}", dist_m.value);
    println!("Feet: {}", dist_f.value);
}

Meters and Feet are empty structs — they hold nothing, exist only as names, and are never instantiated. PhantomData is likewise zero-sized, so Distance<Meters> occupies exactly the same eight bytes as a bare f64. Nothing here costs anything at runtime.

What you buy with it is that Distance<Meters> and Distance<Feet> are different types. A function that takes Distance<Meters> will reject a Distance<Feet> at compile time, and adding the two together won’t typecheck unless you write a conversion. The mistake that cost NASA an orbiter becomes a compiler error.

The same trick shows up wherever you want states or roles enforced statically — Connection<Open> vs. Connection<Closed>, Id<User> vs. Id<Order> over the same underlying u64. The pattern is always: an empty marker struct for each tag, a PhantomData<T> field to make the parameter legal, and the type checker does the rest.

PhantomData has a second job in unsafe code — telling the compiler about ownership and variance it can’t otherwise see, e.g. PhantomData<T> on a raw-pointer struct so the drop checker knows a T is logically owned. That matters when you’re writing your own collection; for the tagging use above, the zero-sized-marker reading is all you need.

Traits

A trait is a set of behaviours a type can implement — closest to an interface in Java or a concept in C++. The part that matters day to day is trait bounds: when a generic function writes T: SomeTrait, it’s saying “this only works for types that can do this particular thing.”

Reading bounds fluently is most of what makes standard library signatures stop looking cryptic.

The traits you’ll actually meet

Open any real Rust codebase and the same few dozen traits account for nearly everything. Worth knowing by name, grouped by what they’re for:

Derivable — you’ll see these in #[derive(...)] constantly

Trait What it gives you
Debug {:?} formatting. Derive it on essentially every type.
Clone Explicit, possibly-expensive duplication via .clone()
Copy Implicit bitwise duplication (no .clone() call needed)
Default T::default() and ..Default::default()
PartialEq / Eq == and !=
PartialOrd / Ord <, >, sort(), BTreeMap keys
Hash HashMap / HashSet keys

Conversion

Trait What it gives you
From<T> / Into<T> Infallible conversion. Implement From, get Into free.
TryFrom<T> / TryInto<T> Fallible conversion, returns Result
FromStr "42".parse::<T>()
AsRef<T> / AsMut<T> Cheap reference-to-reference conversion; why fn open(p: impl AsRef<Path>) accepts &str, String, and PathBuf alike
Borrow<T> Like AsRef but promises Eq/Hash agree — why map.get("k") works on a HashMap<String, V>
ToOwned The &str → String direction; the trait behind Cow

Formatting and errors

Trait What it gives you
Display {} formatting — the human-facing message. Never derivable.
Error Marks a type as an error; enables Box<dyn Error> and ? interop

Iteration

Trait What it gives you
Iterator next(), plus the ~70 adapters that come free with it
IntoIterator What for x in thing desugars to
FromIterator What .collect::<T>() requires of T
Extend .extend(iter) on an existing collection

Ownership, pointers, lifecycle

Trait What it gives you
Drop A destructor. Runs at scope exit; mutually exclusive with Copy.
Deref / DerefMut Smart-pointer transparency and deref coercion (below)
Sized Implicit on every generic parameter; opt out with ?Sized

Concurrency

Trait What it gives you
Send The type can move to another thread
Sync &T can be shared across threads

Both are auto-traits — the compiler implements them for you when all fields qualify. You almost never write these; you read them in error messages.

Operators and callables

Trait What it gives you
Add, Sub, Mul, Neg, … (std::ops) Operator overloading
Index / IndexMut x[i]
Fn / FnMut / FnOnce Closures, in decreasing order of restriction

Ecosystem traits you’ll hit within a week

Serialize / Deserialize (serde), Read / Write / BufRead (std::io), Future (async), Any (runtime downcasting).

Two rules of thumb. Derive Debug on everything, and Clone/PartialEq/Eq/Hash whenever they’re free — derives cost nothing at runtime and unlock APIs you didn’t anticipate needing. And take the trait, not the type, in signatures: impl AsRef<Path> over &Path, impl IntoIterator<Item = T> over Vec<T>, impl Display over String.

Default

Default is the trait for types that know how to produce a sensible zero value — 0 for integers, "" for String, None for Option, an empty Vec. You get it via T::default().

A good way to see why the bound exists at all is Cell::take().

Why does Cell::take() require Default?

Suppose you have:

use std::cell::Cell;

let c = Cell::new(10);

When you call:

let value = c.take();

Rust needs to do two things:

  1. Return the current value (10).
  2. Leave something inside the Cell, because a Cell can never be empty.

So it replaces the old value with T::default(). Conceptually, take() does this:

fn take(&self) -> T
where
    T: Default,
{
    self.replace(T::default())
}

So after:

let c = Cell::new(10);

let x = c.take();

println!("{}", x);       // 10
println!("{}", c.get()); // 0

The 10 is returned, and the cell now contains 0, which is i32::default().

For String:

let c = Cell::new(String::from("hello"));

let s = c.take();

println!("{s}");              // hello
println!("{:?}", c.take());   // ""

The cell is replaced with an empty String.

So whenever you see a bound like:

T: Default

read it as:

“This function only works for types that know how to create a default value using T::default().”

PartialEq and Eq — one method, two promises

== desugars to exactly one thing: PartialEq::eq(&a, &b). That’s the trait with the actual code in it. For a struct whose fields are all comparable, derive it and you get field-by-field comparison:

#[derive(Debug, PartialEq)]
struct User {
    id: u64,
    name: String,
}

Eq, by contrast, is this in full:

pub trait Eq: PartialEq<Self> {}

An empty marker trait. It defines no methods and changes nothing about what runs at the call site — a == b still calls PartialEq::eq, whether or not Eq is implemented. What Eq does is let other APIs demand a stronger promise as a bound. HashMap, HashSet, BTreeMap and Ord all require Eq, because their internals would misbehave if a key weren’t equal to itself.

The three axioms. PartialEq asks for two:

  • Symmetry — if a == b then b == a
  • Transitivity — if a == b and b == c then a == c

Eq asks for those plus a third:

  • Reflexivity — a == a holds for every value, no exceptions

That third one is the entire difference, and floats are the reason it’s carved out. f64::NAN == f64::NAN is false by IEEE 754 design: NaN means “this result is meaningless,” and two meaningless results aren’t equal. Note precisely which axiom that breaks — symmetry is fine (NaN == NaN is consistently false in both directions) and transitivity is fine (no chain of equalities leads to a contradiction, since NaN is equal to nothing at all). Only reflexivity fails. That single broken guarantee is why f32/f64 implement PartialEq but not Eq, and it propagates: any struct with a float field can derive PartialEq but not Eq.

The one line that demonstrates it. If you only remember one thing from this section, make it this:

let x = f64::NAN;

assert!(x != x);        // passes — the only value in Rust not equal to itself
assert_eq!(x == x, false);

That’s the whole justification for splitting the trait in two. Every other numeric type is reflexive, which is why they’re all Eq:

assert!(i32::MAX == i32::MAX);   // true, always
assert!(0u8 == 0u8);

Now watch the derives track that distinction exactly. This compiles:

#[derive(PartialEq)]
struct Reading {
    value: f64,
}

Add Eq and it doesn’t:

#[derive(PartialEq, Eq)]
struct Reading {
    value: f64,
}
error[E0277]: the trait bound `f64: Eq` is not satisfied
   --> eq_fail.rs:3:5
    |
  1 | #[derive(PartialEq, Eq)]
    |                     -- in this derive macro expansion
  2 | struct Reading {
  3 |     value: f64,
    |     ^^^^^^^^^^ the trait `Eq` is not implemented for `f64`
    |
note: required by a bound in `std::cmp::AssertParamIsEq`

That AssertParamIsEq in the last line is the derive machinery showing its hand: #[derive(Eq)] generates no comparison code at all — it just emits a static assertion that every field is itself Eq. Which is the perfect illustration of what Eq is: not behaviour, just a checked claim.

Swap the field for a u64 and both derives work:

#[derive(PartialEq, Eq)]      // fine — u64 and String are both Eq
struct User {
    id: u64,
    name: String,
}

So the rule to state in an interview: derive Eq whenever every field is Eq — it’s free and unlocks HashMap keys — and the one thing that will stop you is a float.

Integers have no such value — every bit pattern of an i32 is an ordinary number, i32::MAX == i32::MAX is true — so they implement both. And integer division by zero isn’t a NaN analogue; there’s no bit pattern to put there, so Rust panics instead. Among floats, only 0.0 / 0.0 (and friends like (-1.0).sqrt(), inf - inf) gives NaN — 1.0 / 0.0 is a well-defined inf.

So should you derive both? For User above, yes — all its fields are Eq, so it’s free, costs nothing at runtime, and unlocks using User as a HashSet element or HashMap key. The only reason to derive PartialEq alone is a field that genuinely isn’t reflexive, i.e. a float. Deriving Eq there would be a lie about your type’s semantics, and the compiler stops you anyway.

The contract is unenforced

Here’s the part worth internalizing: symmetry, transitivity and reflexivity are promises you make, not properties the compiler checks. #[derive(PartialEq)] always produces a well-behaved impl. The moment you hand-write impl PartialEq, nothing stops you from breaking the rules — you just get code that misbehaves at a distance, in sorting or hashing or dedup logic that assumed the contract held.

Case 1: a == b but b != a. A derived impl can’t do this; it’s symmetric by construction. A buggy manual one can:

struct Fuzzy(i32);

impl PartialEq for Fuzzy {
    fn eq(&self, other: &Self) -> bool {
        // bug: the trailing clause makes this directional
        (self.0 - other.0).abs() <= 2 && self.0 > other.0
    }
}

The realistic version of this mistake involves cross-type comparison. PartialEq is generic over the right-hand side — PartialEq<Rhs = Self> — so impl PartialEq<B> for A is legal and gives you a == b. But b == a is then a different impl, resolved separately, and keeping the two consistent is entirely on you:

struct Celsius(f64);
struct Fahrenheit(f64);

impl PartialEq<Fahrenheit> for Celsius {
    fn eq(&self, other: &Fahrenheit) -> bool {
        self.0 == (other.0 - 32.0) / 1.8
    }
}

impl PartialEq<Celsius> for Fahrenheit {
    fn eq(&self, other: &Celsius) -> bool {
        self.0 * 1.8 + 32.0 == other.0   // different arithmetic, different rounding
    }
}

Both impls look right. They compute the conversion in opposite directions, so floating-point rounding can make c == f and f == c disagree for the same pair of values. If you write a cross-type PartialEq, write both directions and make them the same computation.

Case 2: symmetric but not transitive. This is the classic, and it’s why “approximate equality” is a trap. Define equality as “within a tolerance”:

struct Approx(f64);

impl PartialEq for Approx {
    fn eq(&self, other: &Self) -> bool {
        (self.0 - other.0).abs() < 1.0
    }
}

let a = Approx(1.0);
let b = Approx(1.9);
let c = Approx(2.8);

a == b   // true  (diff = 0.9)
b == c   // true  (diff = 0.9)
a == c   // false (diff = 1.8)  ← transitivity broken

Symmetry holds perfectly — (x - y).abs() doesn’t care about argument order. But equality chains drift: a is close to b, b is close to c, and a is not close to c. Feed this type to sort_by or use it as a HashMap key and you get results that depend on comparison order.

This is exactly why f64’s own PartialEq uses exact bit comparison rather than a tolerance. Exact equality keeps symmetry and transitivity intact; it gives up only reflexivity, and only for NaN. If you need tolerance comparison, write it as a named method — fn approx_eq(&self, other: &Self) -> bool — not as PartialEq. Callers then know they’re getting a fuzzy relation instead of assuming the axioms.

PartialOrd and Ord

The same split, one level up. PartialOrd::partial_cmp returns Option<Ordering> — None means “these two aren’t comparable.” Ord::cmp returns a plain Ordering and promises a total order: every pair compares, and the ordering is transitive and antisymmetric.

Floats are the same culprit for the same reason. Here the Option in the return type stops being an abstraction and becomes something you can print:

assert_eq!(1.0_f64.partial_cmp(&2.0), Some(Ordering::Less));   // comparable
assert_eq!(f64::NAN.partial_cmp(&1.0), None);                  // not comparable

None is the whole reason PartialOrd exists. And it has a consequence people find genuinely surprising the first time — with NaN involved, a comparison and its opposite are both false:

let x = f64::NAN;

assert_eq!(x < 1.0, false);
assert_eq!(x > 1.0, false);
assert_eq!(x == 1.0, false);   // all three at once

In every other type, !(a < b) && !(a == b) implies a > b. That’s trichotomy, and it’s precisely what a total order guarantees and a partial one doesn’t. So f64 is PartialOrd but not Ord — which is the concrete cause of the error you hit the first time you sort floats:

let mut v = vec![3.0_f64, 1.0, 2.0];
v.sort();
error[E0277]: the trait bound `f64: Ord` is not satisfied
   --> sortfail.rs:3:7
    |
  3 |     v.sort();
    |       ^^^^ the trait `Ord` is not implemented for `f64`

sort needs Ord because a sorting algorithm has to be able to order any two elements it’s handed — a None mid-sort would leave it with nowhere to go. Three ways out, in increasing order of how much I’d recommend them:

v.sort_by(|a, b| a.partial_cmp(b).unwrap());   // panics the moment a NaN appears
v.sort_by(|a, b| a.total_cmp(b));              // total order, NaN sorts to the end
let mut v = vec![3.0_f64, f64::NAN, 1.0, 2.0];
v.sort_by(|a, b| a.total_cmp(b));
// [1.0, 2.0, 3.0, NaN]

total_cmp is the right default. It implements IEEE 754’s total ordering, so it never panics and gives NaN a defined position instead of pretending it can’t occur.

When you derive PartialOrd/Ord on a struct, comparison is lexicographic in declaration order — first field, then second as a tiebreak, and so on. Reordering fields silently changes sort behaviour, which is a nice trap to know about before it bites you.

Copy and Clone — why String isn’t Copy

Copy means “duplicating this value is just a memcpy of its bytes, and both copies are independently valid.”

A String is three words on the stack — pointer, length, capacity — pointing at a heap allocation. Bitwise-copying those three words gives you two Strings pointing at the same buffer. Both would run their destructor at scope end → double free.

Rust enforces this structurally: Copy and Drop are mutually exclusive. If a type has a destructor, it can’t be Copy, because Copy implies duplication is a no-op the compiler can do silently and the value has no cleanup obligation to duplicate.

struct Foo;
impl Copy for Foo {}
impl Drop for Foo { fn drop(&mut self) {} }
// error[E0184]: the trait `Copy` cannot be implemented for this type;
//               the type has a destructor

So anything owning a resource — String, Vec, Box, File, Rc (needs a refcount increment) — is Clone instead. Clone is the explicit, possibly-expensive version: you have to write .clone(), which makes the allocation visible in the source.

Copy is the “trivially duplicable” subset: integers, char, bool, &T, and aggregates of those.

Rule of thumb:

Copy Clone
Trigger implicit — assignment, function arguments explicit .clone()
Cost always cheap (bitwise copy) can be expensive (deep copy, allocation)
Original still valid after? yes only if you called .clone() — a plain move still invalidates the original
Requires all fields Copy, no Drop impl fields need Clone, or a custom impl
Relationship Copy: Clone — every Copy type must also implement Clone Clone stands alone

The first row is the one to say out loud: Clone you have to ask for, Copy just happens. Cloning is always a visible .clone() in the source — the compiler will never insert one for you, which is deliberate, since a hidden deep copy is exactly the kind of cost you want spelled out. Copy, by contrast, fires silently on ordinary assignment and function calls; there’s no syntax for it at all, because there’s nothing to be careful about.

let a = 5;
let b = a;              // Copy — implicit, nothing written
let s = String::from("hi");
let t = s.clone();      // Clone — explicit, you asked for the allocation

That’s the trade the two traits encode: cheap things get to be invisible, expensive things have to be typed out.

That last row has a practical consequence: Copy is a supertrait of Clone, so you can never derive one without the other.

#[derive(Copy)]
struct P { x: i32 }
error[E0277]: the trait bound `P: Clone` is not satisfied
    |
  1 | #[derive(Copy)]
    |          ---- in this derive macro expansion
  2 | struct P { x: i32 }
    |        ^ the trait `Clone` is not implemented for `P`
    |
note: required by a bound in `Copy`

Which is why you always see the pair written together:

#[derive(Clone, Copy)]
struct P { x: i32 }

For a Copy type the two do exactly the same thing — .clone() on a Copy type is just the bitwise copy, so the derived Clone is *self. Clippy will tell you as much if you write .clone() on one (clippy::clone_on_copy).

The row worth reading twice is the third. Copy doesn’t mean “this value can’t be moved” — it means moves are copies, so the original stays usable:

let a = 5;
let b = a;
println!("{a}");        // fine — i32 is Copy

let s = String::from("hi");
let t = s;
println!("{s}");        // error[E0382]: borrow of moved value: `s`

Same syntax, entirely different semantics, decided solely by whether the type is Copy. That’s the single most common source of early borrow-checker confusion — the move is invisible in the source, and only the type tells you it happened.

Debug vs. Display

Two formatting traits, and the split is about audience.

Debug is {:?} — for programmers. It’s derivable, it’s allowed to be ugly, and it should exist on essentially every public type you write. {:#?} is the pretty-printed multi-line version, which is what you want when inspecting nested structures.

Display is {} — for humans. It is deliberately not derivable, because there’s no way for a macro to guess what a good user-facing message looks like. You write it by hand:

use std::fmt;

impl fmt::Display for User {
    fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
        write!(f, "{} (#{})", self.name, self.id)
    }
}

Implementing Display also gets you .to_string() for free, via a blanket impl<T: Display> ToString for T in the standard library. That’s the same one-impl-many-uses pattern as From/Into.

The rule for error types: implement both. Debug for the developer reading a stack trace, Display for the message the user sees. std::error::Error requires both as supertraits precisely for that reason.

Why Vec<i32> doesn’t implement Display

Sooner or later you write println!("{}", v) on a vector and get:

error[E0277]: `Vec<i32>` doesn't implement `std::fmt::Display`
  = help: the trait `std::fmt::Display` is not implemented for `Vec<i32>`
  = note: in format strings you may be able to use `{:?}` instead

This isn’t an oversight. Display is reserved for types that have exactly one natural, unambiguous, user-facing textual form — numbers, strings, chars, paths. A Vec<i32> is a container, not a value, and there are several defensible ways to print one:

[1, 2, 3]
1, 2, 3
(1, 2, 3)
1 2 3
one per line

The standard library can’t know which you meant, so it declines to choose. Debug has no such problem: its audience is a programmer at a debugger, [1, 2, 3] is a fine answer, and so Vec<T>: Debug holds whenever T: Debug.

The same reasoning is why HashMap, HashSet, BTreeMap, tuples, arrays, slices and Option are all Debug but not Display — they’re structures, not values.

What matters is why the library’s refusal is final, and that’s the orphan rule: if std had picked a format and you disagreed, you could not override it, because you’re not allowed to write impl Display for Vec<i32> in your own crate. Declining to implement it leaves the decision with you.

Three ways forward:

Use Debug and move on. Almost always the right call for logs and inspection.

let v = vec![1, 2, 3];
println!("{:?}", v);    // [1, 2, 3]
println!("{:#?}", v);   // one element per line, indented

Format at the call site when you want a specific separator:

let v = vec![1, 2, 3];
let s = v.iter().map(|x| x.to_string()).collect::<Vec<_>>().join(", ");
println!("{s}");        // 1, 2, 3

Wrap it in a newtype when the same formatting shows up in more than one place. The wrapper is a type you own, so you’re allowed to implement Display for it:

use std::fmt;

struct CommaSeparated(Vec<i32>);

impl fmt::Display for CommaSeparated {
    fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
        for (i, n) in self.0.iter().enumerate() {
            if i > 0 {
                write!(f, ", ")?;
            }
            write!(f, "{n}")?;
        }
        Ok(())
    }
}

println!("{}", CommaSeparated(vec![1, 2, 3]));   // 1, 2, 3

Writing into the formatter directly, rather than building an intermediate Vec<String> and joining it, avoids allocating — write! appends straight to the output buffer. It also means {:>20} and friends behave sensibly if the caller asks for padding.

The orphan rule

You may write impl Trait for Type only if the trait or the type is defined in your own crate. If both come from elsewhere — another crate, or the standard library — the impl is rejected:

error[E0117]: only traits defined in the current crate can be implemented
              for types defined outside of the crate

The four cases:

Trait Type Allowed
local local yes
local foreign yes — implement your trait for Vec<i32> freely
foreign local yes — the everyday impl Display for MyType
foreign foreign no — this is the orphan rule biting

So impl Display for Vec<i32> fails on both counts: Display lives in std, and so does Vec. Neither half is yours.

Why the rule exists: coherence. Trait impls in Rust are global. There’s no importing an impl, no scoping one to a module, no choosing between them at the call site — for a given trait and a given type, the whole program sees exactly one implementation. That’s what lets println!("{}", x) resolve without you naming anything, and what lets HashMap<K, V> trust that every holder of a K hashes it the same way.

Now suppose two crates in your dependency graph both did this:

// crate `pretty`
impl Display for Vec<i32> { /* prints [1, 2, 3] */ }

// crate `csv_ish`
impl Display for Vec<i32> { /* prints 1,2,3   */ }

Your program depends on both. Which one runs at println!("{}", v)? There is no principled answer — and worse, the answer would change based on which crates happen to be in the build, so adding an unrelated dependency could silently change your output or break compilation somewhere deep in the tree. An impl like that, disconnected from both its trait’s crate and its type’s crate, is an orphan instance; forbidding them is what keeps coherence decidable. (Haskell permits orphans with a warning and the ecosystem has been paying for it ever since; Rust chose the strict rule.)

The rule also protects the other direction. Because only std can implement std traits for std types, std can add impl Display for Vec<i32> in a future release without breaking anyone. If downstream crates could have written that impl, every new std impl would be a breaking change.

The newtype escape hatch. Wrap the foreign type in a tuple struct you own:

struct CommaSeparated(Vec<i32>);          // local type
impl fmt::Display for CommaSeparated { }  // foreign trait, local type — allowed

This is the newtype pattern, and it’s the standard answer whenever the orphan rule blocks you. It costs nothing at runtime — a single-field tuple struct has the same layout as the field — but it does cost ergonomics: CommaSeparated has none of Vec’s methods. Implement Deref<Target = Vec<i32>> on it if you want them back, though be aware that using Deref purely for inheritance-like method reuse is a mild anti-pattern (see Deref coercion).

The rule is a little more generous than “one side must be local.” The precise form allows a foreign trait if a local type appears among its generic parameters, before any type parameter that’s fully generic. In practice this means:

struct MyType;

impl From<MyType> for Vec<i32> { … }   // ok: MyType is local and appears first
impl Display for Vec<MyType> { … }     // ok: the local type is inside the foreign one
impl Display for Vec<i32> { … }        // rejected: nothing local anywhere

That second one is the useful loophole: Vec<MyType> counts as sufficiently “yours” because the impl can’t collide with anyone else’s — no other crate knows about MyType. The exact formulation lives in RFC 2451, and it’s worth knowing it exists so you don’t reflexively reach for a newtype when you don’t need one.

Where you’ll hit this in practice: implementing serde::Serialize for a type from a third-party crate. Both foreign, both out of reach. Serde’s answer is #[serde(remote = "…")], which generates a local shadow type — the newtype pattern with the boilerplate written for you. The general shape of the workaround is always the same: make one side local.

Hash — and its contract with Eq

To use a type as a HashMap key or a HashSet element, you need Hash + Eq — and since Eq: PartialEq, that’s three derives in practice:

#[derive(Hash, PartialEq, Eq)]
struct User {
    id: u64,
    name: String,
}

Why both? Because hashing alone can’t identify a key. A hash narrows the search to one bucket; equality confirms you found the right key inside that bucket. Collisions aren’t an edge case to engineer around — they’re guaranteed by pigeonhole, since a u64 hash has to represent infinitely many possible User values. So lookup is always two steps: hash to find the bucket, then == against the candidates in it. Drop Eq and the second step has nothing to call.

The contract between them:

If a == b, then hash(a) must equal hash(b).

Break it and HashMap silently loses entries — you insert a key, look it up with an equal key, and get None, because the two hashed into different buckets and the equality check never even runs. Deriving Hash alongside PartialEq/Eq keeps them in sync automatically. Hand-writing one but deriving the other is how this goes wrong: if your manual PartialEq ignores a field, your Hash must ignore it too.

Note the direction — the reverse isn’t required. Two unequal values may hash the same; that’s just a collision, and == sorts it out.

So is #[derive(Hash, PartialEq)] enough? It compiles fine on its own — but the type still won’t work as a key. The bound lives on HashMap’s methods, not on the derive:

impl<K: Eq + Hash, V> HashMap<K, V> { ... }

so the error lands at the insert/get call site rather than on the type, and it’s an E0599 (“method exists, but its trait bounds were not satisfied”) instead of the E0277 you might expect:

error[E0599]: the method `insert` exists for struct `HashMap<K, &str>`,
              but its trait bounds were not satisfied
  |
5 |     struct K(u64, f64);
  |     -------- doesn't satisfy `K: Eq`
7 |     m.insert(K(1, f64::NAN), "data");
  |       ^^^^^^
  |
  = note: the following trait bounds were not satisfied:
          `K: Eq`
help: consider annotating `K` with `#[derive(Eq, PartialEq)]`

Worth recognising that shape — “method exists but trait bounds were not satisfied” almost always means a missing derive, not a missing method. Always the three-derive version.

Why does HashMap insist on Eq rather than settling for PartialEq? This is the question that makes the whole Eq marker trait earn its keep. PartialEq doesn’t guarantee reflexivity, and a key that isn’t equal to itself breaks the map at a basic level:

#[derive(Hash, PartialEq)]     // no Eq — id is f64
struct Reading { id: f64 }

The compiler stops you there, so to actually watch it break you have to lie — hand-write the impls and claim Eq anyway:

#[derive(Debug, Clone, Copy)]
struct Reading { value: f64 }

impl PartialEq for Reading {
    fn eq(&self, other: &Self) -> bool { self.value == other.value }  // NaN != NaN
}
impl Eq for Reading {}                       // the lie: this type is not reflexive
impl Hash for Reading {
    fn hash<H: Hasher>(&self, h: &mut H) { self.value.to_bits().hash(h); }
}

Note the Hash impl is impeccable — it hashes the raw bits, so two NaNs hash identically. The contract “a == b implies equal hashes” is upheld. Only reflexivity is broken. Now insert one key and try to use it:

let key = Reading { value: f64::NAN };
let mut m = HashMap::new();
m.insert(key, "sensor-7");
len            : 1
get(&key)      : None
contains_key   : false
remove(&key)   : None

The entry is in the map — len says so — and there is no way to reach it. The lookup hashes to the correct bucket, finds the byte-identical key sitting right there, runs == on it, gets false, and reports the key absent. It can’t be read, can’t be removed, and inserting the “same” key again doesn’t replace it:

m.insert(key, "sensor-9");
// len after 2nd  : 2
// [(Reading { value: NaN }, "sensor-7"), (Reading { value: NaN }, "sensor-9")]

Two entries with visually identical keys, both unreachable, growing without bound. Eq is exactly the promise that this can’t happen: every key is findable by an equal key, including itself. The bound isn’t bureaucracy — it’s the map refusing to accept keys it could lose.

BTreeMap and BTreeSet want Ord instead of Hash + Eq, for the same underlying reason one level up — they need a total order to navigate the tree, and PartialOrd returning None would leave a comparison with nowhere to go.

One last thing about that drill snippet, which is a borrow-checker trap rather than a trait one:

map.insert(user, "some data");
let value = map.get(&user);      // error: borrow of moved value

insert takes the key by value, so user is moved into the map and can’t be used afterwards. Fix it by looking up a fresh equal key, deriving Clone and inserting user.clone(), or keying the map on something cheap like user.id. The last option is usually the right one — small Copy keys with the full struct as the value is the more common shape.

From, Into, TryFrom, TryInto — the conversion protocol

The point of these four traits isn’t that they do anything clever. Converting a u64 into a wrapper struct is one line of code with or without them. The point is standardisation.

Without a shared convention, every type invents its own name for the same operation:

UserId::from_u64(42)
UserId::new_from_u64(42)
UserId::parse_u64(42)
UserId::convert(42)

Four crates, four spellings, and you look up the docs every time. Rust’s answer is: there is one name, and everybody uses it.

From — “this conversion is always valid”

#[derive(Debug, Clone, Copy, PartialEq)]
struct UserId(u64);

impl From<u64> for UserId {
    fn from(x: u64) -> Self {
        UserId(x)
    }
}

That single impl gives every Rust programmer two forms they already know how to read:

let a = UserId::from(42);
let b: UserId = 42u64.into();   // same thing, direction driven by the annotation
assert_eq!(a, b);

Note there is no Result anywhere. Implementing From is a claim: every value of the source type maps to a valid value of the target. If that isn’t true, you want TryFrom.

Into — the same conversion, read from the caller’s side

You never implement Into. The standard library has a blanket impl:

impl<T, U> Into<U> for T where U: From<T> { ... }

so From gets you Into free — and writing Into yourself is a hard error, not a style preference:

error[E0119]: conflicting implementations of trait `Into<UserId>` for type `u64`
  = note: conflicting implementation in crate `core`:
          - impl<T, U> Into<U> for T where U: From<T>;

What Into is for is bounds. It reads naturally in argument position, where the source type is the generic one:

fn greet(id: impl Into<UserId>) {
    let id: UserId = id.into();
    println!("hello, user {}", id.0);
}

greet(7u64);          // hello, user 7
greet(UserId(9));     // hello, user 9

The second call works because of another blanket impl — impl<T> From<T> for T, the reflexive one — so UserId: Into<UserId> holds trivially. That’s what makes impl Into<T> arguments pleasant: callers who already have the right type aren’t punished for it.

TryFrom — “this conversion can fail”

Now suppose the target has a validity constraint:

#[derive(Debug, PartialEq)]
struct Port(u16);

#[derive(Debug)]
struct PortOutOfRange(u32);

impl TryFrom<u32> for Port {
    type Error = PortOutOfRange;

    fn try_from(x: u32) -> Result<Self, Self::Error> {
        if x == 0 || x > u16::MAX as u32 {
            Err(PortOutOfRange(x))
        } else {
            Ok(Port(x as u16))
        }
    }
}
println!("{:?}", Port::try_from(8080u32));    // Ok(Port(8080))
println!("{:?}", Port::try_from(70000u32));   // Err(PortOutOfRange(70000))

let p: Result<Port, _> = 443u32.try_into();   // Ok(Port(443))

The Result is the whole message. A caller reading Port::try_from(value)? knows without reading any docs that failure is a real possibility they have to deal with — and one they can deal with, because type Error is theirs to define, not a stringly-typed afterthought.

TryInto is to TryFrom what Into is to From: same blanket-impl relationship, same caller-side ergonomics, used the same way in bounds.

The standard library eats its own cooking here. Widening integer conversions are From (i64::from(7u32) — always fine); narrowing ones are TryFrom:

u8::try_from(200i64)   // Ok(200)
u8::try_from(300i64)   // Err(TryFromIntError(PosOverflow))

That pairing is the cleanest illustration of the whole distinction: same conceptual operation, different trait, because one can lose information and the other can’t.

Where this actually pays off

Three places, roughly in increasing order of how much you’ll care.

Generic code doesn’t need to name the source type. fn greet(id: impl Into<UserId>) accepts u64, UserId, and anything anyone adds a From impl for later — including in a downstream crate you’ve never heard of. The function doesn’t change.

Libraries compose without adapters. If crate A defines UserId with From<u64>, and crate C’s API takes impl Into<UserId>, then crate B handing out plain u64s works with C without either of them knowing the other exists. Nobody writes a glue function per type pair, which is the combinatorial version of the problem the naming convention solves individually.

? is built on From. This is the one you hit first in practice, usually without realising it’s the same machinery:

#[derive(Debug)]
enum ConfigError {
    NotANumber(ParseIntError),
    OutOfRange(TryFromIntError),
}

impl From<ParseIntError> for ConfigError {
    fn from(e: ParseIntError) -> Self { ConfigError::NotANumber(e) }
}
impl From<TryFromIntError> for ConfigError {
    fn from(e: TryFromIntError) -> Self { ConfigError::OutOfRange(e) }
}

fn parse_port(s: &str) -> Result<u16, ConfigError> {
    let n: u32 = s.parse()?;      // ParseIntError  -> ConfigError
    let p = u16::try_from(n)?;    // TryFromIntError -> ConfigError
    Ok(p)
}
parse_port("8080")   -> Ok(8080)
parse_port("http")   -> Err(NotANumber(ParseIntError { kind: InvalidDigit }))
parse_port("70000")  -> Err(OutOfRange(TryFromIntError(PosOverflow)))

Two From impls, and ? silently bridges two foreign error types into one local enum. See Converting error types for more on that.

Summary

Infallible Fallible
Implement this From<T> TryFrom<T> (+ type Error)
Get this free Into<T> TryInto<T>
Call site U::from(x) / x.into() U::try_from(x)? / x.try_into()?
Bound to write impl Into<U> impl TryInto<U>

Rules of thumb

  • Implement From/TryFrom, bound on Into/TryInto. The impl side is where the concrete types live; the bound side is where you want to stay generic.
  • From is a promise of totality. If some inputs are invalid, TryFrom — don’t panic inside a From impl.
  • Every From impl gives you TryFrom too, with Error = Infallible. So a try_from call compiling doesn’t mean the conversion can actually fail.
  • The orphan rule still applies: impl From<Celsius> for f64 is fine (your type is in the impl), impl From<u64> for String is not (both foreign). Newtype when you’re stuck.
  • Bare .into() with nothing to infer from is an error, not a guess — E0283, “type annotations needed”. Annotate the binding or use From explicitly.
  • For &str specifically, prefer FromStr over TryFrom<&str> — it’s what .parse() dispatches to.

The one-sentence version: From/Into are the standard protocol for infallible conversions and TryFrom/TryInto for fallible ones, so that generic code and unrelated libraries can compose without anyone inventing bespoke conversion methods.

Iterator vs. IntoIterator

These two get conflated constantly, but they answer different questions. Iterator is the thing that produces values — it has a cursor and you can ask it for the next one:

pub trait Iterator {
    type Item;
    fn next(&mut self) -> Option<Self::Item>;
}

IntoIterator is the thing that can become an iterator. It’s what makes a type loopable:

pub trait IntoIterator {
    type Item;
    type IntoIter: Iterator<Item = Self::Item>;
    fn into_iter(self) -> Self::IntoIter;
}

for x in thing is pure sugar for IntoIterator::into_iter(thing) followed by repeated next() calls. Every Iterator is also an IntoIterator (it returns itself), which is why you can write for x in v.iter() as well as for x in v.

The part that actually trips people up is that Vec<T>, &Vec<T> and &mut Vec<T> each implement IntoIterator differently, so what x is depends on what you handed the loop.

1. for x in numbers — consumes

for x in numbers {
    // x: i32
}

The iterator takes ownership of the collection:

Vec<i32>
   ↓ IntoIterator
Item = i32

Afterward numbers is no longer usable — it was moved into the loop.

2. for x in numbers.iter() — borrows

for x in numbers.iter() {
    // x: &i32
}
numbers
   ↓
&Vec<i32>
   ↓
Iterator<Item = &i32>

The collection stays usable after the loop. You can read each element but not mutate it through that reference.

3. for x in numbers.iter_mut() — borrows mutably

for x in numbers.iter_mut() {
    *x += 10;
}
numbers
   ↓
&mut Vec<i32>
   ↓
Iterator<Item = &mut i32>

Nothing is consumed, but you get mutable references to the elements, so you can write through them.

The clean mental model:

numbers              ownership
    │
    ├── iter()       → &T
    │
    └── iter_mut()   → &mut T

&numbers and &mut numbers in a for loop are equivalent to numbers.iter() and numbers.iter_mut() respectively — the reference impls of IntoIterator just forward to them. And on the signature side, taking impl IntoIterator<Item = T> instead of Vec<T> lets callers pass a vec, an array, a range, or any adapter chain, which is why it shows up in the “take the trait, not the type” rule above.

copied() and cloned()

Since iter() hands you &T, you often need to get back to T. That’s all copied() does:

let x: u8 = nums.iter().copied().sum();

nums.iter() yields &u8 because it borrows rather than consuming. copied() dereferences each item and copies it, so the chain reads Iterator<Item = &u8> → Iterator<Item = u8> → sum::<u8>(), with the sum’s type inferred from the let x: u8 annotation. It’s exactly .map(|&x| x), and it requires T: Copy. When the element type is only Clone, use cloned() instead — same idea, potentially expensive.

Two footnotes. You could drop the copied() here entirely: the standard library implements Sum<&u8> for u8, so summing references works directly. copied() earns its place when later adapters need owned values, or simply when you’d rather read u8 than &u8 in the chain. And watch the width — sum::<u8>() panics in debug and wraps in release once the total passes 255. Nothing widens on your behalf; sum into u32 if the input might be large.

Associated types

Iterator’s type Item; has shown up twice already without explanation, so here’s the mechanism on its own.

A trait is a list of items an implementor must supply. Usually those are functions, but they don’t have to be — a trait can also demand a constant, or a type:

trait Parser {
    type Output;

    fn parse(&self, input: &str) -> Result<Self::Output, ParseError>;
}

The trait doesn’t say what Output is. It says only that every parser has one, and that parse returns it. Each impl fills in the blank:

struct JsonParser;
struct CsvParser;

impl Parser for JsonParser {
    type Output = JsonValue;
    fn parse(&self, input: &str) -> Result<JsonValue, ParseError> { /* ... */ }
}

impl Parser for CsvParser {
    type Output = Vec<Record>;
    fn parse(&self, input: &str) -> Result<Vec<Record>, ParseError> { /* ... */ }
}

Self::Output in the trait definition is a placeholder that resolves per impl. Forget the type Output = ...; line in an impl and the compiler tells you exactly that: not all trait items implemented, missing: Output.

Naming one from outside

Inside the trait you write Self::Output. Outside, you write it against whatever type you have in hand:

fn run<P: Parser>(p: &P, s: &str) -> Result<P::Output, ParseError> {
    p.parse(s)
}

P::Output is the whole point of the feature. run doesn’t know or care what the parser produces, but it can still name the type in its own signature and hand it back to the caller concretely — call run(&JsonParser, s) and you get a Result<JsonValue, ParseError>, not some opaque thing you have to unwrap through a second type parameter.

When P::Output is ambiguous — P implements two traits that both have an Output — fall back to the qualified form from the UFCS section, which works for types exactly as it does for methods:

<P as Parser>::Output

Putting a bound on it

You can constrain the associated type where you declare it. This says whatever you pick for Item, it has to be cloneable and printable:

trait Container {
    type Item: Clone + std::fmt::Debug;

    fn get(&self, i: usize) -> Option<Self::Item>;
}

An impl that sets type Item to something without those bounds is rejected right there, at the impl — not later, at some distant call site. And the payoff is that generic code over Container gets the bounds for free:

fn print_first<C: Container>(c: &C) {
    if let Some(item) = c.get(0) {
        println!("{item:?}");    // works — Item: Debug is guaranteed by the trait
    }
}

Note what print_first did not have to write: no where C::Item: Debug. The bound is part of the trait’s contract, promised once by the trait and relied on by every consumer, rather than something each function re-states. That’s the trade — you narrow who can implement Container in order to widen what you can do with one.

Or bound it later, on the method that needs it

A trait-level bound like type Item: Clone + Debug is a demand on everyone. Sometimes only one method needs the guarantee, and making it a condition of implementing the trait at all is too strong. Put the bound on that method instead:

use std::fmt::Display;

trait Shape {
    type Unit;

    fn area(&self) -> Self::Unit;

    fn describe(&self) -> String
    where
        Self::Unit: Display,
    {
        format!("Area is {}", self.area())
    }
}

describe is a default method — it comes free with the trait — and it doesn’t know or care what Unit concretely is. It only needs it printable, and it says so locally rather than taxing the trait’s declaration.

struct Circle { radius: f64 }

impl Shape for Circle {
    type Unit = f64;

    fn area(&self) -> Self::Unit {
        3.14 * self.radius * self.radius
    }
}

struct PixelRectangle { width: i32, height: i32 }

impl Shape for PixelRectangle {
    type Unit = i32;

    fn area(&self) -> Self::Unit {
        self.width * self.height
    }
}

fn main() {
    let circle = Circle { radius: 5.0 };
    let rectangle = PixelRectangle { width: 10, height: 20 };

    println!("{}", circle.describe());       // Area is 78.5
    println!("{}", rectangle.describe());    // Area is 200
}

Neither impl mentions Display. Both get describe anyway, because f64 and i32 happen to satisfy the clause at the call site.

The difference from the trait-level bound is worth holding onto, because it’s a difference in when the check happens. A bound on the declaration is checked at the impl: pick a Unit that isn’t Display and the impl itself fails. A where clause on a method is checked at the call: the impl compiles fine, and you simply can’t call describe on it. So a Shape whose Unit is some opaque non-printable measurement is still a perfectly good Shape — it just doesn’t get that one method.

That’s the general shape of the choice. Bound the declaration when the guarantee is part of what the trait means; bound the method when it’s an extra one method happens to want.

Pinning it at the use site: Trait<Assoc = Type>

This is the syntax the post has been using for a while — Iterator<Item = String>, Add<Output = T> — and it’s worth being precise about, because it looks like a generic argument and isn’t one.

fn print_all<I: Iterator<Item = String>>(iter: I) {
    for s in iter {
        println!("{s}");
    }
}

Quiz: how do you write a function bound requiring I to be an Iterator whose Item is String? The correct answer is fn foo<I: Iterator<Item = String>>(iter: I), because Item is an associated type rather than a generic parameter of the trait

If you’ve done any Rust flashcards you’ve probably met this one, and the three wrong answers are wrong in three instructive ways — passing String positionally to a trait that has no positional slot, applying it to I as though I were the generic thing, and using a bare trait name where a type belongs.

Iterator has no generic parameters. The Item = String inside the angle brackets is an equality constraint: it doesn’t select a variant of the trait, it narrows the bound to those implementors whose Item happens to be String. Compare the two roles side by side:

fn a<T: From<i32>>(x: T) {}                    // From<i32> — i32 is an input, part of
                                               // the trait's identity
fn b<I: Iterator<Item = String>>(i: I) {}      // Item = String — a constraint on the
                                               // one Iterator impl I already has

Both live inside <>, which is why they get conflated. The = is the tell. And a constraint like this is not the only way to use the trait — fn c<I: Iterator>(i: I) accepts every iterator; adding Item = String just narrows the set.

Since Rust 1.79 you can also bound the associated type inline, without naming it exactly:

fn shout<I: Iterator<Item: Display>>(iter: I) { /* ... */ }

which before then had to be spelled out longhand as a second clause:

fn shout<I>(iter: I)
where
    I: Iterator,
    I::Item: Display,
{ /* ... */ }

The longhand still works and is often clearer once you have more than one such bound.

Chaining: bounds on an associated type’s associated type

Once you can name I::Item, you can bound it — and the thing you bound it with may itself have an associated type, which you can constrain in the same breath:

fn sum_squares<I>(iter: I) -> i32
where
    I: Iterator,
    I::Item: std::ops::Mul<Output = i32> + Copy,
{
    iter.map(|x| x * x).sum()
}

Read the second clause left to right. I::Item reaches through I’s Iterator impl to name what it yields. That type must implement Mul, and Mul’s own associated Output must be i32 — which is what makes x * x an i32 and lets sum() add them into the declared return type. Copy is there because x * x uses x twice.

So sum_squares accepts an iterator of i32, or of any custom Meters type that multiplies into a plain i32, and rejects an iterator of String — all without a single concrete item type in the signature. This is the everyday reason associated types are worth the trouble: each one is a name you can hang further requirements on, and they nest as deep as you need.

More than one

Nothing says a trait gets only one. A trait describing a graph needs two blanks filled in, because a graph is defined by both what its nodes are and what its edges are:

trait Graph {
    type Node;
    type Edge;

    fn edges(&self, node: &Self::Node) -> Vec<Self::Edge>;
}

struct CityMap;

impl Graph for CityMap {
    type Node = String;
    type Edge = (String, String, u32);   // (from, to, distance)

    fn edges(&self, node: &String) -> Vec<(String, String, u32)> {
        vec![(node.clone(), "NextCity".to_string(), 42)]
    }
}

Both are outputs of the same choice: pick CityMap and you’ve picked cities-as-strings and distance-weighted edges together. Had these been generic parameters, Graph<String, (String, String, u32)> and Graph<u32, (u32, u32, f64)> would be different traits that CityMap could implement both of, and fn shortest_path<G: Graph>(g: &G) would need to thread two extra type parameters through every signature to say what it’s walking over. With associated types it just writes G::Node and G::Edge.

Notice too that the impl writes the concrete types in edges’ signature — &String, not &Self::Node. Either spelling compiles; once type Node = String; is fixed, they’re the same type. Writing them out is usually clearer in an impl, and Self::Node is usually clearer in the trait.

Both mechanisms in one trait: Add

The standard library’s Add is the cleanest real example of a generic parameter and an associated type living side by side, each doing the job the other can’t:

trait Add<Rhs = Self> {
    type Output;

    fn add(self, rhs: Rhs) -> Self::Output;
}

And in use:

use std::ops::Add;

#[derive(Debug, Clone, Copy)]
struct Point { x: f64, y: f64 }

impl Add for Point {
    type Output = Point;

    fn add(self, rhs: Point) -> Point {
        Point { x: self.x + rhs.x, y: self.y + rhs.y }
    }
}

impl Add<f64> for Point {
    type Output = Point;

    fn add(self, scalar: f64) -> Point {
        Point { x: self.x + scalar, y: self.y + scalar }
    }
}

Two impls of Add for one type, and they don’t collide — because Rhs is a generic parameter, part of the trait’s identity, so Add<Point> and Add<f64> are genuinely different traits. That’s what you want: adding a point to a point and adding a scalar to a point are both reasonable, and a type should be allowed to do both.

Output is an associated type for the opposite reason. Once you’ve fixed which impl you’re in — Point + Point, or Point + f64 — the result type isn’t a free choice any more. There’s one sensible answer per impl, and each impl states it once.

So read the declaration as two different questions:

trait Add<Rhs = Self> {
             │
             └── input: what may I be added to?     several answers → generic parameter

    type Output;
         │
         └── output: what falls out?                one answer per impl → associated type

The = Self is a default for the parameter, which is why impl Add for Point works without writing impl Add<Point> for Point. That’s covered in its own section, along with where the Add<Output = T> bound comes from — and Output = T there is exactly the equality-constraint syntax from above, applied to this trait.

Trait objects have to spell it out

dyn Trait is a type, and a type has to be fully known. So a trait object over a trait with an associated type won’t compile until you say what it is:

let it: Box<dyn Iterator> = ...;                  // error: the value of the associated
                                                  //        type `Item` must be specified
let it: Box<dyn Iterator<Item = u32>> = ...;      // fine

The reason is the same one that makes P::Output useful in the generic case, viewed from the other end. A caller holding a Box<dyn Iterator<Item = u32>> needs to know that next() returns Option<u32> — that’s in the type, not in the vtable. Erase the concrete parser or iterator and you can still erase which one it is; you can’t erase what it produces.

Two things it can’t (yet) do

An associated type can’t have a default in the trait, the way a method can:

trait Parser {
    type Output = String;    // error: associated type defaults are unstable
}

That’s the associated_type_defaults feature, still nightly-only. If you want a common case to be free, the usual workaround is a blanket impl or a second trait.

The other limit was lifted in Rust 1.65: an associated type can now take its own generic parameters and lifetimes, which is what generic associated types means.

trait Container {
    type Iter<'a>: Iterator<Item = &'a Self::Item>
    where
        Self: 'a;

    type Item;

    fn iter<'a>(&'a self) -> Self::Iter<'a>;
}

That declares a family of types rather than one, indexed by the lifetime. It’s what lets a trait return a borrowing iterator tied to &self — impossible before, because the associated type had to be one fixed type with no way to mention the caller’s lifetime. GATs get thorny fast and are worth reaching for only when you hit that exact wall.

Why Item is an associated type and not a generic parameter

Iterator could plausibly have been declared like this:

trait Iterator<T> {
    fn next(&mut self) -> Option<T>;
}

It wasn’t, and the reason is worth understanding because the same decision comes up in your own traits.

A generic parameter on a trait is an input: it’s part of the trait’s identity, so Iterator<u8> and Iterator<String> are two different traits as far as coherence is concerned, and one type may implement both.

struct Foo;
impl Iterator<u8> for Foo { /* ... */ }
impl Iterator<String> for Foo { /* ... */ }

That compiles under the generic formulation — the impls don’t overlap. But now Foo is an iterator of u8 and an iterator of String simultaneously, which isn’t what iteration means. A cursor over a sequence produces one kind of thing.

The fallout lands on every piece of generic code that mentions the trait. You’d need a second type parameter just to name the item:

fn total<I: Iterator<T>, T>(iter: I) -> T { /* ... */ }

and, for a type with several impls, inference can’t pick one — you’d be reaching for a turbofish. Multiply that by every combinator in the library. map, filter, zip, chain and flat_map would each thread an extra parameter through their signatures and their return types.

An associated type is an output instead: it’s determined by the implementing type, so Iterator can only be implemented once per type and I::Item is unambiguous.

fn total<I: Iterator>(iter: I) -> I::Item { /* ... */ }

One parameter, no turbofish, and Item falls out of inference everywhere.

The rule generalizes cleanly:

Shape of the relationship Use Examples
One impl per type, output follows from the type Associated type Iterator::Item, Add::Output, Deref::Target
Several impls per type, keyed on varying input Generic parameter From<T>, PartialEq<Rhs>, Add<Rhs>

From is the instructive contrast: you genuinely want impl From<i32> for MyType and impl From<String> for MyType side by side, so the source type has to be a trait input. Add uses both at once — Rhs is a parameter (you can add a Meters to a Meters and to an f64), while Output is associated (once the operands are fixed, the result type isn’t a free choice).

So the associated type isn’t just an ergonomic tidy-up for Iterator. It’s the encoding of a real constraint: a type is an iterator over exactly one thing.

Operator overloading and the Add<Output = T> bound

Overloading + for your own type means implementing std::ops::Add. For a concrete type it’s unremarkable, but doing it generically is where a bound most people write wrong for the first time or two shows up.

The goal is to add two Point<T> values componentwise, for any T that can be added:

use std::ops::Add;

struct Point<T> {
    x: T,
    y: T,
}

impl<T> Add for Point<T> {
    type Output = Point<T>;

    fn add(self, rhs: Point<T>) -> Point<T> {
        Point { x: self.x + rhs.x, y: self.y + rhs.y }
    }
}

That doesn’t compile. T is completely unconstrained, so self.x + rhs.x is asking a type that promised nothing to support +. The obvious fix is to say it supports addition:

impl<T: Add> Add for Point<T> {

Still no. T: Add says T + T is a valid expression, but the trait’s signature is fn add(self, rhs: Rhs) -> Self::Output — the result is T::Output, which the bound leaves completely open. It could be T, it could be anything. The field you’re assigning into is a T, so the compiler has no reason to believe those match.

The bound that works pins down the output too:

impl<T: Add<Output = T>> Add for Point<T> {
    type Output = Point<T>;

    fn add(self, rhs: Point<T>) -> Point<T> {
        Point { x: self.x + rhs.x, y: self.y + rhs.y }
    }
}

T: Add<Output = T> is two promises in one: T + T typechecks, and what comes back is a T. That’s exactly what Point { x: self.x + rhs.x, ... } needs.

The arc in one table:

Bound Problem
impl<T> T might not support + at all
impl<T: Add> T supports +, but the result type is unconstrained — it might not be T
impl<T: Add<Output = T>> T supports + and the result is guaranteed to be T, matching the field type

This is why associated-type bounds — the Trait<Assoc = Concrete> shape — turn up constantly in generic numeric and operator code. Naming the trait is usually not enough; you also have to pin down what it produces. The same pattern reads I: Iterator<Item = u32> when you need an iterator of a specific element type, or T: Deref<Target = str> for anything that derefs to a string slice.

Default type parameters and Rhs = Self

Here’s the full declaration of Add:

trait Add<Rhs = Self> {
    type Output;
    fn add(self, rhs: Rhs) -> Self::Output;
}

Two things in that line are worth unpacking.

Rhs = Self is a default type parameter. Omit the argument and Rhs becomes Self, the implementing type — which is why impl Add for Point<T> (with no <...> after Add) gets rhs: Point<T> for free. The default encodes the common case: you usually add a thing to another thing of the same kind. Overriding it is how you get the exceptional case, like scaling a point by a scalar:

impl Mul<f64> for Point<f64> {
    type Output = Point<f64>;

    fn mul(self, scalar: f64) -> Point<f64> {
        Point { x: self.x * scalar, y: self.y * scalar }
    }
}

Self::Output is qualified because Output is a member of the trait, not a free-standing name in scope. Writing bare Output would be like writing x where you meant self.x. The qualification says “the Output belonging to whatever type is implementing this.”

You can use Self as a default in your own traits — a default value can be any type expression that makes sense there: a concrete type, another parameter in scope, or Self:

trait Combine<Other = Self> {
    fn combine(self, other: Other) -> Self;
}

// Other defaults to Self = Point<T>
impl<T: Add<Output = T>> Combine for Point<T> {
    fn combine(self, other: Point<T>) -> Self {
        Point { x: self.x + other.x, y: self.y + other.y }
    }
}

Which brings the two mechanisms in Add into one picture. Rhs is a generic parameter because a type may legitimately be addable to several different things, and each needs its own impl — Add<Point> and Add<f64> for the same Point are distinct traits and don’t overlap. Output is an associated type because once both operands are fixed, the result type isn’t a free choice any more; there’s exactly one right answer per impl. Multiplicity where you want choice, a single fixed value where you don’t — the reasoning is the same as Iterator::Item’s, applied to both slots of one trait.

Fn, FnMut, FnOnce

Closures aren’t a single type. Every closure you write gets its own anonymous struct holding whatever it captured, and the three Fn* traits describe how calling it touches that captured state. The difference is entirely in how each takes self:

Trait Call signature Meaning
Fn fn call(&self, ...) callable through a shared reference; call it as many times as you like, concurrently
FnMut fn call_mut(&mut self, ...) needs exclusive access; callable repeatedly, one at a time
FnOnce fn call_once(self, ...) consumes the closure; callable exactly once

They nest: every Fn is also a FnMut, and every FnMut is also a FnOnce. So FnOnce is the loosest bound (accepts the most closures) and Fn is the strictest.

How the compiler decides

You never annotate which trait a closure implements — Rust infers it from the closure body:

Closure body does… Implements
Nothing with the environment, or only reads via & Fn (+ FnMut, FnOnce)
Mutates a captured variable FnMut (+ FnOnce)
Moves a captured variable out FnOnce only
let name = String::from("world");

let greet = || println!("hello, {name}");   // reads only        → Fn
greet();
greet();                                     // fine, repeatable

let mut count = 0;
let mut bump = || count += 1;                // mutates capture   → FnMut
bump();
bump();

let consume = || drop(name);                 // moves `name` out  → FnOnce
consume();
// consume();                                // ❌ use of moved value

move is a separate axis. The move keyword forces the closure to capture by value rather than by reference — it doesn’t decide which trait you get. A move closure that only reads its captures is still Fn:

let data = vec![1, 2, 3];
let show = move || println!("{}", data.len());  // owns `data`, still Fn
show();
show();

You reach for move when the closure has to outlive the scope that created it — spawning a thread, returning a closure, storing one in a struct.

Which bound to write. Take the loosest trait your function actually needs, since that accepts the most callers:

  • Calling it once (Option::map, thread::spawn) → FnOnce
  • Calling it repeatedly with state (Iterator::for_each, retry loops) → FnMut
  • Calling it repeatedly, possibly from several threads (Iterator::map, callbacks) → Fn
fn call_twice<F: Fn()>(f: F)      { f(); f(); }
fn tally<F: FnMut()>(mut f: F)    { f(); f(); }   // note: `mut f`
fn run<F: FnOnce()>(f: F)         { f(); }

Two practical notes. A FnMut parameter has to be bound as mut f — you can’t call call_mut through an immutable binding. And function pointers (fn(i32) -> i32) implement all three, so a plain fn item can be passed anywhere a closure bound is expected.

Where each one shows up in practice

Iterator adapters — the most common place you’ll meet them:

// closures that only read: these are Fn
let doubled: Vec<i32> = v.iter().map(|x| x * 2).collect();
let evens: Vec<&i32>  = v.iter().filter(|x| **x % 2 == 0).collect();

// closures that carry state: these are FnMut
let mut total = 0;
v.iter().for_each(|x| total += x);

let mut seen = HashSet::new();
let deduped: Vec<&i32> = v.iter().filter(|x| seen.insert(**x)).collect();

Note the bound on all of these — map, filter, for_each, find, any, all are declared as FnMut, even though most closures you pass them only read. That’s deliberate: FnMut is the looser bound, so Fn closures satisfy it automatically and stateful ones work too.

Sorting and comparators — same story. sort_by and sort_by_key are also declared FnMut, though the comparators you actually write are almost always pure Fn:

v.sort_by(|a, b| a.cmp(b));
v.sort_by_key(|x| x.priority);

Option / Result combinators — these take FnOnce, because the closure runs at most once:

let y = x.unwrap_or_else(|| expensive_default());   // FnOnce, and only if x is None
let mapped = r.map_err(|e| format!("wrapped: {e}")); // FnOnce

map, and_then, or_else, unwrap_or_else, ok_or_else — all FnOnce.

Spawning threads and tasks — also FnOnce, since the closure is invoked exactly once when the task runs:

std::thread::spawn(move || {
    do_work(data);
});

thread::spawn and tokio::spawn both want FnOnce() -> T + Send + 'static.

Callbacks and event handlers — stored in a struct, so they need a Box<dyn ...>. Which trait depends on whether the hook can fire more than once:

struct Button {
    on_click: Box<dyn FnMut()>,          // fires repeatedly, may track a click count
}

struct OneShotTimer {
    on_fire: Option<Box<dyn FnOnce()>>,  // fires once, then consumed
}

The Option on the one-shot is load-bearing: calling a Box<dyn FnOnce()> moves it, so you need .take() to get it out of the struct.

Retry loops and polling — FnMut, since the closure is called repeatedly and usually tracks something across attempts:

fn with_retry<F, T, E>(mut attempt: F, tries: u32) -> Result<T, E>
where
    F: FnMut() -> Result<T, E>,
{
    for _ in 0..tries.saturating_sub(1) {
        if let Ok(v) = attempt() { return Ok(v); }
    }
    attempt()
}

Your own generic APIs — the case that matters once you’re writing library-ish code:

fn process<F>(items: &[Item], mut handler: F)
where
    F: FnMut(&Item) -> bool,
{
    for item in items {
        if handler(item) { break; }
    }
}

A workable heuristic when you’re unsure: start at Fn, loosen to FnMut when the compiler complains that a capture needs mutating, and loosen to FnOnce only when the closure genuinely consumes something. Each step outward accepts strictly more callers, so you want to land on the loosest bound your call pattern actually permits.

Supertraits: naming a pile of bounds

Write a generic numeric function and the bound list gets out of hand fast. Something that needs arithmetic, casting, Display, iterator sums and a known maximum ends up dragging ten traits behind every signature. Since a trait can inherit from other traits, you can name that pile once:

pub trait Number:
    Num
    + FromPrimitive
    + ToPrimitive
    + Copy
    + Display
    + Debug
    + Sum
    + Product
    + AddAssign
    + SubAssign
    + MulAssign
    + DivAssign
    + Bounded
    + NumCast
{}

The empty body is the point — Number adds no methods of its own. Everything after the colon is a supertrait: to implement Number a type must already implement all of them, and in return every generic function that writes T: Number gets the whole set.

Most of those come from the num crate. Num covers the arithmetic operators, FromPrimitive / ToPrimitive / NumCast handle generic numeric conversion, and Bounded supplies min_value() and max_value(). Sum and Product are from std::iter — they’re what make .sum() and .product() work on an iterator of T.

The payoff is that signatures go back to being readable:

fn mean<T: Number>(xs: &[T]) -> T {
    xs.iter().copied().sum::<T>() / T::from_usize(xs.len()).unwrap()
}

There are two ways to hand out the impls. Explicitly, one line per type:

impl Number for i32 {}
impl Number for f64 {}
// ...

or with a blanket impl covering anything that qualifies:

impl<T> Number for T
where
    T: Num + FromPrimitive + ToPrimitive + Copy + /* ...the rest... */
{}

The blanket version is less typing and picks up user-defined numeric types for free — someone’s fixed-point or big-integer type becomes a Number the moment it satisfies the bounds. The cost is that you’ve given up control of the membership: every qualifying type is a Number whether you meant it or not, and because the blanket impl already covers them, nobody — including you — can write a manual impl Number for MyType afterwards without a coherence error. Listing the primitives by hand keeps the set deliberate and closed, which is usually what a library wants.

Rust’s aliasing rules

The next few sections lean on this constantly, so it’s worth pinning down.

Aliasing simply means multiple references pointing at the same memory.

let mut x = 10;

let a = &x;
let b = &x;
        x
        │
      +----+
      | 10 |
      +----+
      ▲    ▲
      │    │
      a    b

That’s aliasing. Rust’s fundamental rule about it is:

Many immutable references, or one mutable reference. Never both at the same time.

✔ &T, &T, &T, &T
✔ &mut T
✘ &T and &mut T simultaneously
✘ two &mut T simultaneously

Often summarised as exclusive mutation, shared reading.

Why the rule exists

Suppose Rust allowed this:

let mut x = 5;

let a = &x;
let b = &mut x;

println!("{}", *a);
*b = 100;
println!("{}", *a);

What should a see — 5 then 100, or 5 then 5, or something else?

The compiler wants freedom to optimise. When it sees let a = &x; it assumes nobody can mutate x while a exists, so it’s free to cache *a in a register:

register = 5

println(register);

*b = 100;

println(register);

That prints 5 and 5, even though memory now holds 100. The optimisation is only correct because the language guarantees the mutation can’t happen. So the rule is really a promise to the optimiser: if an immutable reference exists, nobody may mutate the value. (This is the noalias guarantee that comes up again in the UnsafeCell section below.)

Two mutable references are worse still:

let mut x = 0;

let a = &mut x;
let b = &mut x;

*a += 1;
*b += 1;

Which one runs first? Rust refuses to answer the question — it just says there can be only one mutable reference.

Everything in the next section is about bending this rule without breaking it. RefCell, in particular, enforces exactly the same rule; it just checks it while the program is running instead of while it compiles.

Interior mutability

Rust’s default deal is simple: &T means read-only. If all you hold is a shared reference you can’t change what’s behind it, and the optimizer is allowed to assume nobody else will either.

Sometimes that’s too strict. Cell<T>, RefCell<T> and OnceCell<T> all let you mutate through a &T anyway. All three wrap UnsafeCell underneath, all three are !Sync. The difference is what each one does to keep aliasing honest.

The problem, concretely

Here’s the smallest program that hits the wall:

struct Counter {
    count: i32,
}

impl Counter {
    fn increment(&self) {
        self.count += 1;
    }
}
error[E0594]: cannot assign to `self.count`, which is behind a `&` reference

The obvious fix is &mut self, and most of the time that’s the right answer. But sometimes you can’t have it. Maybe the object is already shared:

let config = Config::new();
let a = &config;
let b = &config;

Two live shared references means nobody can take a &mut, ever — that’s the aliasing rule, not a technicality. And yet perhaps the only thing that needs to change is a cache-hit counter buried in one field. Or you’re implementing a trait whose method signature takes &self and you don’t get a vote. Or you’re inside a Fn closure rather than an FnMut one.

Cell is what you reach for then:

use std::cell::Cell;

struct Counter {
    count: Cell<i32>,
}

impl Counter {
    fn increment(&self) {
        self.count.set(self.count.get() + 1);
    }

    fn value(&self) -> i32 {
        self.count.get()
    }
}

let c = Counter { count: Cell::new(0) };
c.increment();
c.increment();
assert_eq!(c.value(), 2);

Still &self, and the counter moves. The mental model: &T normally means read-only, and a Cell field carves out one small mutable box inside an otherwise shared object.

Cell: move values in and out

Cell<T> never hands out a reference to its interior. No reference means nothing to alias, so no checking is needed. Zero runtime cost, can’t panic.

let c = Cell::new(5);
c.set(10);              // no &mut needed
let v = c.get();        // requires T: Copy
let old = c.replace(7); // works for non-Copy too
let owned = c.take();   // requires T: Default

The Copy bound on get() looks arbitrary until you ask what Cell::<String>::get() would do. The String stays in the cell and you also get one back: two owners of one heap allocation, two eventual frees. So get() is restricted to types where copying the bytes really is the whole story. Non-Copy types are still perfectly good cell contents — they just have to leave rather than be duplicated, via replace, take or swap.

The other catch is no in-place mutation. To modify a Cell<Vec<i32>> you take it out, push, and set it back.

RefCell: real references, checked at runtime

RefCell<T> keeps a borrow counter. borrow() gives a Ref<T>, borrow_mut() gives a RefMut<T>, and the count is decremented when those guards drop.

let rc = RefCell::new(vec![1, 2, 3]);
rc.borrow_mut().push(4);
println!("{}", rc.borrow().len());

let _a = rc.borrow();
let _b = rc.borrow_mut();  // panics: already borrowed

Conceptually the counter walks through three states. It starts unborrowed; each borrow() bumps it to shared-with-n-readers; borrow_mut() is only allowed from unborrowed and moves it to exclusive. Ask for an exclusive borrow while two readers are alive and there’s no legal answer, so it panics:

thread 'main' panicked at 'already borrowed: BorrowMutError'

Costs a word of storage plus a branch per borrow, and violations panic at runtime instead of failing to compile. try_borrow / try_borrow_mut return a Result if you’d rather handle it than fall over.

Wait — isn’t borrow checking a compile-time thing?

It is, for real references. let a = &mut x; let b = &mut x; is a compile error: no runtime state, no cost, no way to get it wrong.

RefCell doesn’t get that, and the reason isn’t that rustc is weak. From the compiler’s point of view RefCell is an ordinary library type:

pub struct RefCell<T> {
    value: UnsafeCell<T>,
    borrow: Cell<isize>,   // roughly
}

borrow() is a normal method call that reads and writes a normal integer field. Nothing in the language says that field means “borrow state” — that meaning lives entirely inside the implementation. The borrow checker reasons about & and &mut, and Rust deliberately keeps it that way; the alternative is a compiler that has to understand the semantics of every library type anyone ever writes.

And even if it did understand RefCell, the question is frequently unanswerable:

let cell = RefCell::new(10);
let a = cell.borrow();
maybe_modify(&cell, user_clicked_button());  // takes borrow_mut() if the flag is set

Whether that panics depends on a button. Same story for a recursive function that only calls borrow_mut() at depth 5, or a borrow_mut() sitting after a loop that may never terminate — deciding “can these two borrows overlap?” in general means deciding whether the loop halts, which Turing showed in 1936 that no algorithm can do. The number of paths isn’t really the obstacle: one branch on runtime input is already enough.

So there are two options. Reject every program you can’t prove safe, and lose a pile of correct ones. Or check while running. RefCell picks the second, and the price is a word of storage, a compare-and-branch per borrow, and a panic where you’d have preferred a compiler error.

Cell RefCell OnceCell
Access get/set whole value &/&mut to interior & after first write
Cost free borrow flag + check one Option check
Failure can’t fail panics on bad borrow second set() returns Err
Good for small Copy types large or non-Copy types write-once / lazy init

Rule of thumb: Cell first for Copy scalars, it’s strictly cheaper and can’t blow up. RefCell when you need to operate on the value in place, which is why Rc<RefCell<T>> is the standard shape for shared mutable graphs.

If you hold a &mut Cell<T> or &mut RefCell<T>, both have get_mut(), which is free and statically checked. The runtime machinery only exists on the shared path.

Across threads: Mutex and RwLock

None of the above is thread-safe, and not by oversight. Picture two threads calling borrow_mut() on the same RefCell: both read the flag as unborrowed, both conclude they’re clear, both write “exclusive”, and now two threads hold a &mut to the same value. That’s a data race — precisely what the whole type system exists to rule out.

Rust’s answer is to make it not compile. Cell and RefCell are !Sync, so they can’t be shared across threads at all:

let x = RefCell::new(5);
thread::spawn(move || {
    println!("{}", *x.borrow());
});
error[E0277]: `RefCell<i32>` cannot be shared between threads safely

The thread-safe versions live in std::sync and are shaped the same way — a guard you hold, and a release that happens when it drops:

use std::sync::{Mutex, RwLock};

let m = Mutex::new(5);
{
    let mut v = m.lock().unwrap();  // guard, much like RefMut
    *v += 1;
}                                   // unlocked here, on drop

let rw = RwLock::new(5);
let r = rw.read().unwrap();         // many readers at once
drop(r);
let w = rw.write().unwrap();        // or one writer, exclusive

Mutex is one-at-a-time. RwLock is borrow/borrow_mut promoted to threads: any number of readers, or a single writer. Reach for Mutex by default and RwLock only when reads genuinely dominate — it’s the heavier lock, and a read-mostly workload can still lose the theoretical win to the extra bookkeeping.

Why not just give RefCell an atomic counter and call it thread-safe? Because detecting the conflict was never the hard part. Make the flag an AtomicIsize, let two threads race for exclusive access, and the atomic will reliably tell the loser it lost. Then what? It can’t proceed, and it can’t reasonably panic — both threads were following the rules. It needs to wait. That’s what a lock adds on top of the atomic: the memory ordering that makes the previous holder’s writes visible when you acquire, and a way to block until your turn. (lock() returns a Result for one reason: if a thread panicked mid-update while holding it, the data may be half-written, so the lock is poisoned and everyone afterwards is told.)

The progression is worth holding in your head:

Checked Cost Failure mode
&T / &mut T compile time none won’t compile
Cell<T> nothing to check none none
RefCell<T> runtime, one thread flag + branch panics
Mutex / RwLock runtime, across threads lock, and blocking blocks; can deadlock

For a lone Copy scalar shared between threads, skip the locks and use an atomic (AtomicUsize and friends) — it plays the role Cell plays single-threaded. And when the shared value needs shared ownership as well, the pairs are Rc<RefCell<T>> on one thread and Arc<Mutex<T>> across several.

OnceCell: write once, then it’s frozen

OnceCell<T> is the third shape, and it trades generality for something the other two can’t do: it hands out a &T that stays valid.

It starts empty and can be filled exactly once. After that the value never moves and never changes, so a plain reference to the interior is safe to give out — there’s no possible mutation to alias against.

use std::cell::OnceCell;

let cell: OnceCell<String> = OnceCell::new();

assert!(cell.get().is_none());          // Option<&String>

cell.set(String::from("hello")).unwrap();   // Ok(()) the first time
assert!(cell.set(String::from("bye")).is_err());  // Err(value) — handed back

let s: &String = cell.get().unwrap();   // a real reference, no guard, no Copy bound

set returns Result<(), T> — on failure you get your value back rather than losing it. The method you’ll actually use most is get_or_init, which does the check-and-fill in one step:

struct Parser {
    source: String,
    tokens: OnceCell<Vec<Token>>,
}

impl Parser {
    fn tokens(&self) -> &[Token] {          // &self, returns a borrow
        self.tokens.get_or_init(|| tokenize(&self.source))
    }
}

Compare that to the Cell<Option<u64>> memoization example further down. Cell forces the cached type to be Copy and makes you return a copy every call; OnceCell caches a Vec and lends it out. That’s the reason to reach for it: lazily initialized fields behind &self, where the cached thing is expensive or non-Copy.

The catch is in the name — once. There’s no set after the first, no invalidating the cache, no recomputation. If the value needs to change, you want Cell or RefCell.

A few relatives worth knowing:

  • OnceLock<T> — the thread-safe version, in std::sync. Same API, blocks other threads racing to initialize. This is what you want for a global (static CONFIG: OnceLock<Config> = OnceLock::new();), since OnceCell is !Sync and can’t be a static.
  • LazyCell<T, F> / LazyLock<T, F> — the initializer is baked in at construction, so it fires on first deref instead of at an explicit get_or_init call. LazyLock is the modern replacement for the lazy_static! and once_cell::sync::Lazy you’ll see in older code.
  • The once_cell crate predates all of these. It’s still common in the wild, but for new code the standard library versions are stable and there’s no reason to take the dependency.

Passing them around

The signature is just a shared reference:

fn solve(counter: &Cell<u64>) {
    counter.set(counter.get() + 1);
}

(ref is a keyword, so don’t name the parameter that.)

There’s no stable update method, so read-modify-write is manual: c.set(c.get() * 2).

Two conversions worth knowing. Cell::from_mut turns a &mut T into a &Cell<T> for free, when you already have exclusive access but want to hand out several shared-mutable views. And as_slice_of_cells turns a &Cell<[T]> into a &[Cell<T>], so you can mutate elements while several parts of the code hold the slice.

Don’t reach for this too early. If a single caller owns the value, plain &mut T is better: compile-time checked, costs nothing. Cell earns its place when you need aliasing and mutation. The classic case is a memoized field on a struct whose methods take &self:

struct Grid {
    data: Vec<i32>,
    checksum: Cell<Option<u64>>,
}

impl Grid {
    fn checksum(&self) -> u64 {   // &self, not &mut self
        if let Some(v) = self.checksum.get() { return v; }
        let v = self.data.iter().map(|&x| x as u64).sum();
        self.checksum.set(Some(v));
        v
    }
}

Callers holding shared references can still call it. That’s the whole payoff.

Cell without get()

The Copy bound is only on get() — see the Copy and Clone section above for why String and friends can’t have it. Cell<T> itself works for any T, and the move-out API covers plenty: take(), replace(), set(), swap(). The pattern is pull the value out, work on the owned value, put it back.

struct Collector {
    pending: Cell<Vec<Event>>,
}

impl Collector {
    fn push(&self, e: Event) {
        let mut v = self.pending.take();  // leaves Vec::new() behind
        v.push(e);
        self.pending.set(v);
    }

    fn drain(&self) -> Vec<Event> {
        self.pending.take()               // hand off, reset to empty
    }
}

The canonical one is single-threaded async:

struct Signal {
    waker: Cell<Option<Waker>>,
}

impl Signal {
    fn register(&self, w: Waker) { self.waker.set(Some(w)); }
    fn fire(&self) {
        if let Some(w) = self.waker.take() { w.wake(); }
    }
}

Also Cell<Option<Box<Node>>> for unlinking nodes in linked structures behind a shared reference, and front.swap(&back) for double-buffering — no Copy, no clone, no allocation.

One footgun with take(): while you hold the value, the cell contains T::default(). Anything that reaches back into that cell mid-operation sees an empty value instead of the real one. RefCell would panic loudly in the same situation, which is sometimes what you want.

For non-Copy types RefCell is usually nicer anyway — self.pending.borrow_mut().push(e) beats the take-push-set dance. Cell<T> for non-Copy shines specifically when the operation is a move: take it, swap it, hand it off.

It’s UnsafeCell all the way down

Every type in this section — Cell, RefCell, OnceCell, and their sync counterparts Mutex, RwLock, OnceLock — is a safe wrapper around the same primitive: UnsafeCell<T>.

That matters because mutating through a shared reference isn’t something you can build yourself out of ordinary Rust. &T carries a hard guarantee to the optimizer that the pointee won’t change, and the compiler emits noalias on that basis. UnsafeCell<T> is the only thing in the language that opts out of it — it’s a compiler-recognized lang item, not a clever library trick.

So the whole hierarchy is one unsafe primitive plus different strategies for keeping the aliasing honest:

Wrapper Strategy
Cell never hand out a reference at all
RefCell count borrows at runtime, panic on conflict
OnceCell allow one write, then it’s immutable
Mutex / RwLock block other threads

Each one takes on the obligation of proving the aliasing rules hold, so your code doesn’t have to. Writing a correct wrapper is genuinely hard — I’ll do a full post on UnsafeCell, noalias, and what the compiler is actually promising soon.

Deref coercion

Write enough Rust and you’ll hit a moment where a function wants a &str, you have a String, and passing &my_string just… works. No conversion, no .as_str(). That’s deref coercion, and it’s worth understanding rather than treating as magic.

fn greet(name: &str) {
    println!("hello, {name}");
}

let owned = String::from("world");
greet(&owned);   // &String, but the function wants &str — this compiles

A useful mental model

Think of deref coercion as Rust saying:

“I expected a reference to type B, but you gave me a reference to type A. If A knows how to dereference into B, I’ll do that automatically.”

The common conversions:

&String   ──►  &str
&Vec<T>   ──►  &[T]
&Box<T>   ──►  &T
&Rc<T>    ──►  &T
&Arc<T>   ──►  &T

No heap allocation or copying happens — it’s simply adjusting the view of the existing data through references.

Why it matters in practice

The payoff is on the API side. If you write a function that takes &str instead of &String, callers can hand you a String, a &str, or a string literal and all three work. Take the narrower type and you only accept one of them:

fn takes_str(s: &str)     { /* accepts &String, &str, and literals */ }
fn takes_string(s: &String) { /* only accepts &String */ }

Same rule for slices: prefer &[T] over &Vec<T> in signatures, since &Vec<T> coerces to &[T] but not the reverse. This is why idiomatic Rust signatures are full of &str and &[T] — the borrowed view is strictly more general than the owned container.

One limit to keep in mind: coercion only goes in the direction Deref defines. &String → &str is free; going the other way costs an allocation and you have to ask for it explicitly with .to_string() or .to_owned().

Autoderef on field and method access

Coercion at a call site is one half of it. The other half shows up whenever you reach into a smart pointer for a field or a method. Here’s the shape that raises the question, straight out of the interior-mutability section:

struct Node {
    next: RefCell<Option<Rc<Node>>>,
}

let a = Rc::new(Node { next: RefCell::new(None) });

How are we able to write a.next here, considering next is a field of the Node struct and that’s one layer deep?

The interesting part is that a is not a Node. Its type is Rc<Node>. So why does this compile?

a.next

instead of

(*a).next

The answer is automatic dereferencing, or autoderef.

Imagine Rc were just a normal struct:

struct Rc<T> {
    value: T,
}

Then you’d have to write (*a).next, because a is an Rc<Node>, *a is a Node, and next is a field of Node.

But the real Rc<T> implements Deref:

impl<T> Deref for Rc<T> {
    type Target = T;
    fn deref(&self) -> &T { /* ... */ }
}

This means Rust knows how to convert Rc<Node> into &Node when needed.

So when the compiler sees a.next, it first asks: does Rc<Node> have a field called next? No. Then it asks: can I dereference it? Yes, because Rc implements Deref. So it effectively becomes (*a).next, where *a uses Rc’s Deref implementation.

a.next
   │
   ▼
(*a).next

This rewriting is automatic.

It works for multiple layers too. With an Rc<Box<Node>>, x.next would become roughly (*(*x)).next. The compiler keeps dereferencing until it finds the field or method you’re trying to access.

The same thing happens with methods. When you write a.next.borrow_mut(), the compiler is doing something like (*a).next.borrow_mut(). You almost never write the explicit * yourself, because Rust inserts these dereferences automatically for field and method access.

So the rule of thumb is: if a type implements Deref, Rust will automatically dereference it for field access (a.next) and method calls (a.foo()), until it finds the requested field or method. This is why smart pointers like Box<T>, Rc<T> and Arc<T> feel almost transparent to use — you can usually interact with the underlying value as if you had it directly.

Autoref vs. autoderef

TL;DR

Autoderef follows Deref to find the method — it looks through a pointer. Autoderef answers “where does this method live?”

Autoref inserts the & or &mut to pass the receiver — it takes a reference to a value. Autoref answers “how do I hand the receiver over?”

A single method call usually does both, in that order: deref down to the type that owns the method, then ref back up to match its self.

Box<Person>  ──autoderef──▶  Person  ──autoref──▶  &Person  ──▶  greet()

And the one-line caveat that explains most surprises: both are method-call features. p.greet() gets them; greet(p) gets neither.

The section above covered autoderef. Autoref is the other half, and it’s the one that’s easier to miss — because what it inserts is a character you never typed.

The & you never wrote

struct Person {
    name: String,
}

impl Person {
    fn greet(&self) {
        println!("Hello {}", self.name);
    }
}

let p = Person { name: "Alice".to_string() };
p.greet();

p is a Person. But fn greet(&self) is shorthand for fn greet(self: &Person) — the method wants a &Person, and you handed it a Person. The types don’t match, and yet it compiles, because the compiler rewrites the call:

p.greet();          // what you write
Person::greet(&p);  // what it resolves to

That inserted &p is autoref.

Why it’s easy to miss: normal functions don’t do this

Here’s the contrast that makes it stick. Same signature, written as a free function:

fn greet(p: &Person) {
    println!("Hello {}", p.name);
}

greet(p);    // error[E0308]: expected `&Person`, found `Person`
greet(&p);   // fine — you write the & yourself

The exact same mismatch that Rust silently fixes in p.greet() is a hard error in greet(p). Autoref is a method-call feature. The compiler does not insert & into arbitrary function calls; it only does it when resolving a receiver.method(args) expression, because that’s the one place it knows what the target wants.

Keep those two lines side by side in your head:

person.greet();    // autoref happens
greet(person);     // it doesn't

It inserts &mut too

Whichever the method asks for:

struct Counter { value: i32 }

impl Counter {
    fn get(&self) -> i32   { self.value }
    fn bump(&mut self)     { self.value += 1; }
}

let mut c = Counter { value: 10 };

c.get();     // → Counter::get(&c)
c.bump();    // → Counter::bump(&mut c)

This is also why c has to be declared mut: the autoref for bump is a &mut borrow, so the usual borrow rules apply to a reference you never wrote. When the compiler says “cannot borrow c as mutable”, the borrow it’s complaining about is often an invisible one.

The two together

Now put a pointer in front of it:

let x = Box::new(Person { name: "Alice".into() });
x.greet();

x is a Box<Person>. greet is defined on Person, and wants &Person. It helps to watch the compiler do this in two distinct passes, because they’re answering two different questions.

Step 1 — method lookup. The compiler sees Box<Person> and goes looking for a greet. There isn’t one on Box<Person>, so it dereferences and looks again:

Box<Person>
     ↓
  Person          // found it: Person::greet

That’s autoderef, and note what it produced — a method, not a receiver. At this point the compiler knows which function it’s calling and nothing else.

Step 2 — determine self. Now it reads that method’s signature. greet takes &self, which here means &Person, and what’s on hand is a Person. So it takes a reference:

  Person
     ↓
 &Person

That’s autoref. Stack the two and you get the whole call:

Box<Person>
    │  autoderef — Box<T>: Deref<Target = T>, so look inside
    ▼
  Person
    │  autoref — greet wants &self
    ▼
 &Person

which is to say Person::greet(&*x). Deref down to find the method, ref back up to call it — and the clean separation is that autoderef is driven by where the method lives, while autoref is driven by what its self asks for. Neither one knows about the other.

It repeats as far as needed. A Box<Box<Person>> derefs twice before the autoref, giving Person::greet(&**x). And it’s why String gets str’s methods free:

let s = String::from("hello");
s.is_empty();          // is_empty is on str, not String
                       // String ──deref──▶ str ──ref──▶ &str

Here’s the same thing once more, stripped to nothing so there’s no name field to distract you:

struct Person;

impl Person {
    fn greet(&self) {
        println!("hello");
    }
}

fn main() {
    let person = Person;
    let boxed = Box::new(person);

    boxed.greet();
}

boxed is a Box<Person>, and the only greet in the program is Person::greet, which wants &Person. So the single line boxed.greet() is doing this:

                    METHOD LOOKUP
                         │
                         ▼
                    Box<Person>
                         │
                     autoderef
                         ▼
                      Person
                         │
                   found greet()
                         │
                         ▼
                 method wants &self
                         │
                      autoref
                         ▼
                     &Person

Writing the * yourself

The useful experiment is to take one of those steps back from the compiler:

(*boxed).greet();

Now you’ve done the dereference by hand, and the picture shortens by a step:

  boxed
    │
    │  *  — you wrote this one
    ▼
  Person
    │
    │  autoref — still the compiler's job
    ▼
 &Person
    │
    ▼
  greet()

The autoref doesn’t go away. You can hand the compiler a Person directly, but you still can’t hand it the &Person that greet actually wants — not through dot syntax, anyway. For that you’d have to drop to Person::greet(&*boxed), which is the fully desugared form with nothing left implicit.

Which is a good moment to separate four words that get used as if they were one thing. * is the syntax for dereferencing explicitly, and it’s the only one of the four you ever type. Deref is the trait that defines what dereferencing means for a custom type — implement it and * starts working on yours. Autoderef is the compiler inserting those derefs on its own while it hunts for a method. Autoref is the compiler inserting the & or &mut once it’s found one, to match whatever the receiver’s self asks for.

Keep autoderef = method lookup and autoref = satisfy the self type in your head and most of this area stops being mysterious.

The actual algorithm

Worth knowing precisely, because it explains the surprises. For receiver.method(args) the compiler:

  1. Builds a candidate list by starting at the receiver’s type and dereferencing repeatedly: Box<Box<Person>>, Box<Person>, Person. (An unsized coercion, e.g. [T; N] to [T], is appended at the end.)
  2. Walks that list in order, and for each candidate type T looks for a method taking T, then &T, then &mut T — in that order.
  3. Takes the first hit.

Two consequences fall out of the ordering.

Earlier in the deref chain wins. The outermost type is checked first, so a method on Box<T> would shadow one on T. This is exactly why the standard library gives Box, Rc and Arc almost no inherent methods, and why the ones they do have are written as associated functions — Rc::clone(&a), not a.clone(). If Rc had an inherent clone method it would shadow T::clone for every T you ever put in one.

&T is tried before &mut T. So a shared borrow is preferred, and you only get a mutable one when nothing else fits.

And the reason to know step 2 at all: it’s the mechanism behind the .clone() trap from the UFCS section. With a &&NotClone receiver, candidate #1 is &NotClone — and &NotClone is Clone, because shared references are always Copy. The search succeeds on the first candidate and stops, so you get a copied reference rather than the error you’d want. Nothing went wrong; the algorithm just found a legitimate match earlier than you expected.

Qualified paths get neither

This is where autoref connects back to UFCS. Take a trait method:

trait Container<T> {
    fn get(&self, index: usize) -> T;
}

impl Container<i32> for MyBox {
    fn get(&self, index: usize) -> i32 { /* ... */ }
}

Called as a method, autoref applies as usual:

my_box.get(0);                                  // → <MyBox as Container<i32>>::get(&my_box, 0)

Written as a qualified path, it does not:

<MyBox as Container<i32>>::get(0);              // error: this function takes 2 arguments
                                                //        but 1 argument was supplied
<MyBox as Container<i32>>::get(&my_box, 0);     // correct

The error is confusing the first time, because you counted one argument and the compiler counted two. The missing one is self — in path form it’s an ordinary first parameter, and there’s no receiver expression sitting to the left of a dot for the compiler to borrow. You pass it yourself, & included.

That’s the rule worth carrying:

Form Autoderef Autoref
x.method(args) yes yes
Type::method(&x, args) no no
<Type as Trait>::method(&x, args) no no
f(x) — a free function no no

Dot syntax is the only place the compiler does this work for you. Every other form is the desugared one, where the borrows are yours to write — which is the same reason a qualified path is a good debugging tool: it shows you what the dot was hiding.

Deref vs. AsRef

Rust gives you two ways to make a type act like another type — but they’re not interchangeable.

Deref is for smart pointers. Implement it, and the compiler implicitly coerces your type wherever a reference to its target is expected — that’s how Box<T>, Rc<T>, and String transparently expose T’s (or str’s) methods without any unwrapping.

AsRef<T> is the explicit, no-magic cousin. It says “call .as_ref() and get a cheap &T view” — no compiler coercion, just a plain conversion. It shines in generic function signatures:

use std::path::Path;

fn open<P: AsRef<Path>>(path: P) {
    let path: &Path = path.as_ref();
    // &str, String, and PathBuf all land here
}

open("config.toml");
open(String::from("config.toml"));
open(std::path::PathBuf::from("config.toml"));

The rule of thumb: reach for Deref only on true single-field wrapper types where implicit coercion is expected. Reach for AsRef whenever you want a flexible, allocation-free API surface. One is compiler magic; the other is explicit convention — know which one you’re signing up for.

Deref AsRef<T>
Who invokes it the compiler, implicitly you, via .as_ref()
Targets per type exactly one many (String: AsRef<str>, AsRef<[u8]>, AsRef<Path>, …)
Typical use smart pointers and newtypes generic parameters
Cost free free

Can a struct have multiple Deref impls?

No — you can only implement Deref once per struct.

impl Deref for MyString {
    type Target = str;
    fn deref(&self) -> &str { &self.inner }
}

impl Deref for MyString {  // ❌ compile error
    type Target = [u8];
    fn deref(&self) -> &[u8] { self.inner.as_bytes() }
}

This fails with conflicting implementations of trait Deref for type MyString.

Why: Deref declares Target as an associated type, not a generic parameter.

pub trait Deref {
    type Target: ?Sized;
    fn deref(&self) -> &Self::Target;
}

Rust’s trait coherence rules say a type can only have one impl of a given trait — and since Target isn’t part of the trait’s identity (it’s just an associated item), the two impl Deref for MyString blocks collide regardless of what Target is set to. The compiler needs *s and method-lookup auto-deref to resolve to exactly one type, so allowing multiple targets would make deref coercion ambiguous by construction.

This is precisely where AsRef<T> differs and wins: T there is a generic type parameter on the trait itself, so AsRef<str> and AsRef<[u8]> are different trait implementations, not conflicting ones:

impl AsRef<str> for MyString {
    fn as_ref(&self) -> &str { &self.inner }
}

impl AsRef<[u8]> for MyString {
    fn as_ref(&self) -> &[u8] { self.inner.as_bytes() }
}

Both compile fine — you just need enough type information at the call site to pick which one applies, either from context or from an annotation:

let s = MyString { inner: String::from("hi") };

let text: &str   = s.as_ref();          // inferred from the annotation
let bytes        = AsRef::<[u8]>::as_ref(&s);  // spelled out explicitly

If you actually need multiple “deref-like views”, the idiomatic pattern is to pick your one Deref::Target to be the most fundamental representation (usually str if you’re wrapping a String-like type) and implement AsRef for any additional views. That mirrors what std itself does: String implements Deref<Target = str> — one canonical target — but AsRef<str>, AsRef<[u8]>, and AsRef<OsStr> all simultaneously.

Box<T>: one owner, on the heap

Now that Deref is in hand, Box is easy to describe: it’s the simplest smart pointer in the language, and *b works on it for exactly the reason the last two sections spelled out.

First, a naming question worth settling. Box is a generic struct in std, not a primitive type — Box<i32> is a type instantiated from it, the same way Vec<i32> is. So “Box<T> is a smart pointer” and “Box is a generic struct” are both right; “Box is a primitive type” isn’t. It does get some help from the compiler, which we’ll get to, but there’s nothing keyword-ish about it.

let x = Box::new(42);
stack                     heap
+---------+              +------+
| pointer | -----------> |  42  |
+---------+              +------+

The value lives on the heap; the Box itself is one pointer wide and lives wherever the variable lives. That’s the entire idea.

It owns what it points to

A Box is not a borrow. It owns the heap value, moves like any other owned value, and frees the allocation when it drops:

let a = Box::new(5);
let b = a;      // ownership moves; `a` is no longer usable
                // heap value freed when `b` goes out of scope

No free, no destructor to remember — Drop runs automatically at the end of the scope. Single owner, which is the distinction from Rc<T> (several owners, one thread) and Arc<T> (several owners, several threads) that showed up in the interior-mutability section.

*b and the auto-deref

Box<T> implements Deref<Target = T> and DerefMut, so everything from the previous two sections applies to it directly:

let x = Box::new(5);
println!("{}", *x);        // 5 — Deref::deref, then the *

let s = Box::new(String::from("hello"));
println!("{}", s.len());   // Box<String> → String → .len()

That’s the general rule worth internalising: *ptr is not a built-in operation on some blessed set of pointer types. Outside of plain references and raw pointers, * on any value is a call to Deref::deref (or DerefMut::deref_mut where a mutation needs it). If a type supports *, someone implemented Deref for it — that’s true of Box, Rc, Arc, Ref, MutexGuard, and every smart pointer you’ll write yourself.

Box does get one privilege the others don’t: you can move the value out through the dereference.

let boxed = Box::new(String::from("hello"));
let owned: String = *boxed;   // moves out, frees the box — only Box can do this

Try that with an Rc and you get “cannot move out of dereference”, because the other owners would be left pointing at nothing. Box knows it’s the only owner, so the compiler special-cases it.

What it’s actually for

Three reasons come up over and over.

Recursive types, the classic one. This doesn’t compile:

enum List {
    Cons(i32, List),   // ❌ recursive type has infinite size
    Nil,
}

To lay out List the compiler has to know how big it is, and that definition says “an i32 plus a List” all the way down. A Box cuts the recursion, because a pointer’s size is known no matter what’s on the other end:

enum List {
    Cons(i32, Box<List>),
    Nil,
}

Same trick for tree nodes: left: Option<Box<Node>>.

Trait objects. A dyn Trait is unsized — the whole point is that the concrete type isn’t known — so it can’t be stored in a variable directly. Putting it behind a pointer gives it a size again:

let items: Vec<Box<dyn Display>> = vec![
    Box::new(5),
    Box::new("hi"),
    Box::new(3.5),
];

Box<dyn Error> in return types is the same pattern, and it’s why Result<T, Box<dyn Error>> is the usual signature for “this can fail and I don’t want to enumerate how” in application code.

Large values you’d rather not move around. Passing a struct by value copies its bytes; passing a Box copies one pointer.

struct Big { data: [u8; 100_000] }

let x = Box::new(Big { data: [0; 100_000] });   // moves are now 8 bytes

Be a little skeptical of this one in the naive form above — Box::new takes its argument by value, so a debug build will typically build the Big on the stack and then copy it to the heap. Release builds usually elide that, but if the value is genuinely enormous the reliable fix is to construct it in place (e.g. build a Vec and box that) rather than trusting the optimizer.

Where the compiler helps

Mostly Box is an ordinary struct you could nearly write yourself, but it isn’t only library code. Beyond the move-out-of-*b privilege above, the compiler knows a boxed recursive type is sized, and it performs unsized coercions on it — Box<String> to Box<dyn Display>, Box<[T; 3]> to Box<[T]> — which no user-defined pointer type gets to do without nightly features.

In one sentence: Box<T> is a generic smart-pointer struct that owns a heap-allocated value, hands it out through Deref, and frees it when it drops.

Rc<T> and Arc<T>: many owners

Box is one owner. When a value genuinely needs several — a node reachable from two places in a graph, a config every part of the program holds on to — you want a reference-counted pointer instead.

use std::rc::Rc;

let a = Rc::new(String::from("hello"));
let b = Rc::clone(&a);      // count is now 2, nothing was copied

println!("{}", Rc::strong_count(&a));   // 2

Rc::clone doesn’t clone the String; it bumps a counter and hands back another pointer to the same allocation. Each Rc that drops decrements the count, and the value is freed when it hits zero. Like Box, Rc<T> implements Deref<Target = T>, so a.len() works and everything from the autoderef section applies.

What you don’t get is mutation. Rc only ever lends out &T — with several owners around, handing out a &mut T would break the aliasing rule outright. That’s why Rc<RefCell<T>> is such a common pairing: Rc for the shared ownership, RefCell for the mutation, each doing one job.

Arc<T> is the same thing with an atomic counter, for when the owners are on different threads. The A is for atomic, not async. It’s slightly slower than Rc because incrementing the count has to be an atomic operation, which is the whole reason both types exist rather than just the safe one.

Why Arc<RefCell<T>> doesn’t compile

This is the trap everyone hits once:

use std::sync::Arc;
use std::cell::RefCell;

let x = Arc::new(RefCell::new(0));
let y = Arc::clone(&x);

std::thread::spawn(move || {
    *y.borrow_mut() += 1;       // ❌ doesn't compile
});

The important thing to remember is that Arc only solves ownership across threads. It does not make the contents thread-safe.

Go through the types one at a time. x is an Arc<RefCell<i32>>. Arc::clone(&x) creates another Arc pointing at the same RefCell. Then thread::spawn moves that Arc<RefCell<i32>> into another thread.

thread::spawn requires that everything captured by the closure is Send, so Arc<RefCell<i32>> has to be Send. And Arc is thread-safe only if its contents are:

impl<T: Send + Sync> Send for Arc<T> {}

Notice it requires T: Sync, not just T: Send. Why? Because multiple threads can dereference the same Arc simultaneously:

Thread A        Thread B

Arc ─┐          Arc ─┐
     ▼               ▼
     RefCell<i32>

Both threads end up holding a &RefCell<i32> to the same cell — and “&T can be shared across threads” is precisely what Sync means.

So: is RefCell<i32> Sync? No. Recall what it stores:

value       = 0
borrow flag = 0

borrow_mut() moves that flag from 0 to -1, and dropping the borrow moves it back. Those updates are ordinary integer reads and writes. Two threads at once:

Thread A                Thread B

read flag = 0
                        read flag = 0
write -1
                        write -1

both think they own the mutable borrow

The borrow flag itself has a data race, and Rust’s aliasing rules are broken — two live &mut to the same i32. RefCell performs no synchronisation whatsoever, so RefCell<T> is !Sync. Therefore Arc<RefCell<i32>> is !Send, and thread::spawn won’t take it:

error[E0277]: `RefCell<i32>` cannot be shared between threads safely
    = help: the trait `Sync` is not implemented for `RefCell<i32>`

The fix is to replace the borrow bookkeeping with real synchronisation:

use std::sync::{Arc, Mutex};

fn main() {
    let x = Arc::new(Mutex::new(0));
    let y = Arc::clone(&x);

    let handle = std::thread::spawn(move || {
        *y.lock().unwrap() += 1;
    });

    handle.join().unwrap();

    println!("{}", *x.lock().unwrap());   // 1
}

Mutex lets only one thread mutate the value at a time, which makes Mutex<T> Sync (given T: Send), which makes Arc<Mutex<T>> safe to share. Arc<RwLock<T>> works the same way when reads dominate.

Type Interior mutability? Thread-safe? Typical use
Cell<T> yes, by moving values in and out no single-threaded, Copy types
RefCell<T> yes, with runtime borrow checks no single-threaded
Rc<T> no — shared ownership only no single-threaded
Arc<T> no — shared ownership only yes, the count only multi-threaded
Mutex<T> yes, via locking yes multi-threaded
RwLock<T> yes, many readers or one writer yes multi-threaded

The common source of confusion is thinking “Arc makes anything inside it thread-safe.” It doesn’t. Arc only makes the reference counting atomic; the safety of the contained value is still T’s responsibility. Rc<RefCell<T>> on one thread, Arc<Mutex<T>> across several.

Rc::make_mut: copy-on-write

Since Rc won’t lend you a &mut T, how do you ever change the value? One answer is Rc::make_mut, which uses a neat trick called copy-on-write. It asks a single question: am I the only owner?

If yes, there’s nobody to disturb, so it just hands you a &mut T:

let mut x = Rc::new(String::from("hello"));   // strong_count == 1

Rc::make_mut(&mut x).push_str(" world");      // no clone, no allocation

If no, mutating in place would be visible to the other owners, so it clones the value first and points your Rc at the fresh copy:

let mut a = Rc::new(String::from("hello"));
let b = Rc::clone(&a);

Rc::make_mut(&mut a).push_str(" world");

println!("{}", a);   // hello world
println!("{}", b);   // hello

Before, both pointed at one allocation:

a ----\
       \
        ▼
     "hello"
        ▲
       /
b -----/

make_mut saw a count of 2, cloned, and left b where it was:

a ---> "hello world"

b ---> "hello"

Why not just always clone? Because you’d pay for it even when nobody was sharing. Written by hand the eager version is:

let mut b = (*a).clone();   // always clones, even at count 1
b.push_str(" world");

Now imagine an Rc<Vec<u8>> holding 500 MB, where 99% of the time there’s a single owner. Cloning first means allocate 500 MB, copy 500 MB, mutate — every time, for nothing. make_mut allocates and copies only in the rare shared case. Hence the name: you pay for the copy only when there’s actually another owner’s view to preserve.

The other half of the appeal is that the caller doesn’t have to know whether the Rc is shared. make_mut means “I want mutable access; if I’m unique don’t waste time copying, and if I’m not, preserve the other owners” — and it works that out itself. Picture an image editor holding Rc<Image>: open the same image in a second tab and an edit forks the pixels, but with one tab open the edit is in place and free.

Two details worth knowing: make_mut needs T: Clone (obviously), and it also clones if any Weak references are outstanding, not just strong ones. If you’d rather not clone at all, Rc::get_mut returns Option<&mut T> — Some when unique, None otherwise. Arc::make_mut and Arc::get_mut are identical.

Rc::try_unwrap: getting the value back out

Rc::try_unwrap is for when you want the owned T back out of the Rc — but only if you’re the last owner.

pub fn try_unwrap(this: Rc<T>) -> Result<T, Rc<T>>

It consumes the Rc and returns Ok(T) if the count was 1, or Err(Rc<T>) if other owners still exist.

Why can’t you just write let s = *x;? Because the allocation might be owned by someone else too:

x ----\
       \
        ▼
     "hello"
        ▲
       /
y -----/

If *x moved the String out, y would be left pointing at nothing. So Rust forbids moving T out of an Rc<T> in general, and try_unwrap is the checked way to ask for it anyway:

let x = Rc::new(String::from("hello"));

let s = Rc::try_unwrap(x).unwrap();   // count was 1 — the String is moved out
println!("{}", s);

No cloning, no allocation: the Rc disappears and the String itself comes back. With a second owner around it fails instead:

let x = Rc::new(String::from("hello"));
let y = Rc::clone(&x);

match Rc::try_unwrap(x) {
    Ok(s)   => println!("got {s}"),
    Err(rc) => println!("still shared: {rc}"),   // this arm
}

The Err arm is why the signature returns Result<T, Rc<T>> rather than Result<T, ()>: try_unwrap consumed your Rc, so on failure it has to hand it back or you’d have lost it.

The usual pattern is building something with shared ownership and then collapsing it at the end, once you know the other references are gone:

let rc = build_graph();
// ... lots of shared ownership ...
drop(other_references);

let graph = Rc::try_unwrap(rc).unwrap();   // now a plain Graph

From there you own a Graph instead of an Rc<Graph> and can move it around with no reference counting at all.

It’s worth being clear about how this differs from make_mut, since both start by asking “am I unique?”:

Method Gives you If there are multiple owners
Rc::make_mut(&mut rc) &mut T clones T (copy-on-write)
Rc::try_unwrap(rc) T returns Err(Rc<T>), no clone

make_mut says “I need mutable access, and cloning to get it is fine.” try_unwrap says “I want the original value — if I can’t have it without cloning, don’t clone, just tell me.” Which makes try_unwrap free when it succeeds: it takes ownership of the existing T without copying or allocating anything.

Weak<T>: a pointer that doesn’t own

The shortest way to say it: Rc<T> owns a value, and Weak<T> points at a value without keeping it alive.

That matters as soon as your data structure has cycles in it — the classic one being a parent that holds its children and children that want to point back at their parent.

use std::rc::{Rc, Weak};

struct Node {
    value: i32,
    parent: Weak<Node>,
    children: Vec<Rc<Node>>,
}

fn main() {
    let parent = Rc::new(Node {
        value: 10,
        parent: Weak::new(),
        children: Vec::new(),
    });

    let child = Rc::new(Node {
        value: 20,
        parent: Rc::downgrade(&parent),
        children: Vec::new(),
    });

    println!("parent strong count = {}", Rc::strong_count(&parent));
    println!("parent weak count   = {}", Rc::weak_count(&parent));

    // Weak doesn't give direct access to the Node.
    // We have to "upgrade" it to an Rc.
    if let Some(parent_rc) = child.parent.upgrade() {
        println!("parent value = {}", parent_rc.value);
    }
}

The line doing the work is Rc::downgrade(&parent). It turns an owning handle:

Rc<Node>
   |
   v
┌───────────┐
│ Node      │
│ value: 10 │
└───────────┘

into a non-owning one:

Weak<Node>
   |
   |  "I know where it is,
   |   but I don't own it."
   v
┌───────────┐
│ Node      │
│ value: 10 │
└───────────┘

So why not just use Rc<Node> for the parent link as well? Suppose we did:

struct Node {
    value: i32,
    parent: Option<Rc<Node>>,
    children: Vec<Rc<Node>>,
}

Now the parent owns the child through children:

parent Rc
   ↓
 child Rc

and the child owns the parent right back through parent:

parent Rc
   ↑
 child Rc

which is a cycle:

parent → child → parent → child → ...

Neither count ever reaches zero, so neither node is ever freed. That’s a leak — safe Rust, no undefined behaviour, memory gone anyway. With Weak on the way up, only one direction counts:

parent
  │
  │ Rc
  ↓
child
  │
  │ Weak
  └────────→ parent

Only the Rc edge keeps anything alive, so dropping the parent actually drops it.

The other thing to know about Weak is that you can’t read through it directly. This doesn’t compile:

child.parent.value

because a Weak<Node> carries no guarantee that the Node is still there. You have to ask first:

if let Some(parent) = child.parent.upgrade() {
    println!("{}", parent.value);
}

upgrade() means “if the object is still alive, give me an Rc to it”, and its return type says exactly that:

Weak<Node>
    │
    │ upgrade()
    ↓
Option<Rc<Node>>
    │
    ├── Some(Rc)  → object still alive
    │
    └── None      → object was dropped

The mental model that makes all of this stick is a short one. Rc<T> is “I own this.” Weak<T> is “I know about this, but I don’t own it.” Which is why Weak is the right tool for parent pointers in trees, graphs, observer relationships, caches, and anywhere else an Rc cycle would otherwise form.

One subtle point worth stating outright: a Weak doesn’t keep anything alive. Once the last Rc goes, the value is dropped even if a pile of Weaks are still pointing at where it used to be — and every upgrade() from then on returns None.

Error propagation and ?

Rust has no exceptions. A function that can fail returns Result<T, E> — either Ok(value) or Err(error) — and the caller has to deal with both arms. Done by hand, that gets verbose fast:

fn read_config() -> Result<String, io::Error> {
    let contents = match fs::read_to_string("config.toml") {
        Ok(c) => c,
        Err(e) => return Err(e),
    };
    Ok(contents)
}

The ? operator collapses that match into one character:

fn read_config() -> Result<String, io::Error> {
    let contents = fs::read_to_string("config.toml")?;
    Ok(contents)
}

? and Ok(...) are inverses

This is the part that clicks late for a lot of people. Ok(...) converts a successful value into a Result. The ? operator does the opposite: it extracts the value from a Result, or returns the error early if there is one.

A nice way to remember it:

?         :  Result<T, E>  →  T          (or early Err)
Ok(...)   :  T             →  Result<T, E>

So in the function above, ? unwraps on the way in and Ok re-wraps on the way out. That’s why almost every fallible function ends with Ok(something) — the return type demands a Result, and you’re handing back a plain value.

Once you see the symmetry, the shape of everyday Rust makes sense:

fn parse_port() -> Result<u16, Box<dyn Error>> {
    let raw = env::var("PORT")?;    // Result<String, VarError> → String
    let port: u16 = raw.parse()?;   // Result<u16, ParseIntError> → u16
    Ok(port)                        // u16 → Result<u16, _>
}

Two different error types flow through that function, and neither is handled explicitly. ? handles them by leaving.

Where ? can be used

? only works inside a function whose return type can absorb the early exit — a Result, an Option, or anything implementing Try. Put it in a function returning () and you get a compile error, which is the single most common first encounter with the operator:

fn main() {
    let contents = fs::read_to_string("config.toml")?;
}
error[E0277]: the `?` operator can only be used in a function that returns
              `Result` or `Option` (or another type that implements `FromResidual`)
 --> src/main.rs:2:52
  |
1 | fn main() {
  | --------- this function should return `Result` or `Option` to accept `?`
2 |     let contents = fs::read_to_string("config.toml")?;
  |                                                     ^ cannot use the `?` operator
  |                                                       in a function that returns `()`

Read the error literally and it tells you the fix: ? needs somewhere to return the Err to, and () has no room for one. The mention of FromResidual is the general version — Result and Option are just the two types in the standard library that implement it.

So the fix is to widen main’s return type — main is allowed to return a Result:

fn main() -> Result<(), Box<dyn Error>> {
    let contents = fs::read_to_string("config.toml")?;
    println!("{contents}");
    Ok(())
}

Note the Ok(()) at the end — the unit value wrapped in Ok, the same T → Result<T, E> move as before, just with nothing interesting inside.

? works on Option too, with the same shape: it unwraps Some(v) to v and returns None early. What it can’t do is mix the two — ? on an Option inside a function returning Result won’t compile. Convert explicitly with .ok_or(...) when you need to cross that boundary.

Converting error types

The other thing ? quietly does is call From::from on the error. If the function returns Result<T, MyError> and you ? on something producing io::Error, it compiles as long as MyError: From<io::Error>:

impl From<io::Error> for MyError {
    fn from(e: io::Error) -> Self {
        MyError::Io(e)
    }
}

That single impl is what lets a function with one error type call into libraries with several. Box<dyn Error> is the low-ceremony version of the same trick — it accepts any error type, at the cost of losing the ability to match on which one you got. Fine for main and small tools; for a library, define a real error enum.

? vs. unwrap()

People agonise over this one, but the rule is actually very simple:

  • Use ? when you want to let the caller handle the error.
  • Use unwrap() when you’re certain failure is impossible (or you deliberately want the program to crash).
// The caller decides what a missing file means.
fn load(path: &str) -> Result<String, io::Error> {
    let contents = fs::read_to_string(path)?;
    Ok(contents)
}

// This regex is a literal I wrote myself; if it doesn't compile that's a bug,
// not a runtime condition to recover from.
let re = Regex::new(r"^\d+$").unwrap();

The dividing line is who has enough context to make a decision. A library function usually doesn’t — it can’t know whether a missing config file is fatal or expected — so it propagates with ? and lets the caller choose. A crash, by contrast, is a claim: this cannot fail, and if it does the program’s assumptions are broken and continuing is worse than stopping.

Where unwrap() is genuinely fine:

  • Tests — a failed assumption should fail the test loudly.
  • Prototypes and one-off scripts, where error handling is noise.
  • Invariants you can prove hold, like parsing a hardcoded literal.

Where it isn’t: anything reading from the network, filesystem, or user. Those fail routinely and a panic just means the failure gets reported with a worse message.

When you do reach for it, prefer expect("...") over bare unwrap(). Same behaviour, but the panic message says what you assumed instead of leaving a stack trace to decode:

let port = env::var("PORT").expect("PORT must be set");
Want Use
Propagate the error to the caller ?
Handle both arms here match
Substitute a default on error .unwrap_or(default)
Crash on error (prototypes, tests) .unwrap() / .expect("msg")
Turn Option into Result .ok_or(err) / .ok_or_else(...)

Getting started with Serde

Serde splits serialization into two concerns, and understanding the split is most of the battle:

  • The serde crate defines the generic Serialize/Deserialize traits, plus a #[derive(...)] macro (enabled via the derive feature) that auto-implements them for your structs.
  • Format-specific crates like serde_json handle the actual encoding and decoding for one format.

So serde describes what your data looks like structurally, and serde_json decides how that gets written to bytes. Swapping JSON for YAML or MessagePack means changing the format crate, not your structs.

Once a type derives these traits, converting to and from JSON is two function calls — serde_json::to_string and serde_json::from_str — both of which return a Result that composes cleanly with anyhow::Result through the ? operator described above.

Basic derive

use serde::{Serialize, Deserialize};

#[derive(Debug, Serialize, Deserialize)]
struct BlogPost {
    id: u32,
    title: String,
}

Deserializing JSON into a struct

let data = r#"{"id": 1, "title": "Hello, Rust"}"#;
let post: BlogPost = serde_json::from_str(data)?;
println!("{:?}", post);

Serializing a struct back to JSON

let json = serde_json::to_string(&post)?;
println!("{}", json);
// {"id":1,"title":"Hello, Rust"}

Pretty-printed JSON

let pretty = serde_json::to_string_pretty(&post)?;
println!("{}", pretty);

Renaming fields to match external JSON conventions

Rust wants snake_case fields; most JSON APIs hand you camelCase. Rather than compromising your struct’s naming, annotate it:

#[derive(Debug, Serialize, Deserialize)]
#[serde(rename_all = "camelCase")]
struct BlogPost {
    id: u32,
    post_title: String, // serializes as "postTitle"
}

Skipping optional/empty fields

#[derive(Debug, Serialize, Deserialize)]
struct BlogPost {
    id: u32,
    title: String,
    #[serde(skip_serializing_if = "Option::is_none")]
    subtitle: Option<String>,
}

Required Cargo.toml setup

[dependencies]
serde = { version = "1", features = ["derive"] }
serde_json = "1"
anyhow = "1"

The gotcha: forgetting the derive feature

This one costs people an afternoon. Leave the derive feature off:

serde = "1"   # missing features = ["derive"]

and the derive macro simply doesn’t exist. The confusing part is what the compiler tells you: it reports that your type doesn’t implement Serialize/Deserialize — even though the #[derive(Serialize, Deserialize)] attribute is sitting right there in the code. The error points at the symptom, not the cause.

Whenever a derive that’s plainly written in your source seems to have had no effect, check the crate’s feature flags first.

Why Rust doesn’t have goroutines

Go’s goroutines are cheap because each one starts with a small stack (a few KB) that grows and shrinks on demand, and Go’s runtime multiplexes many of them onto a few OS threads. Java threads and Rust’s std::thread are the opposite: real OS threads with fixed-size stacks, so memory and context-switching costs cap you in the thousands.

Execution unit OS thread? Stack
Java Thread Yes Fixed (~1 MB, configurable)
Go goroutine No (M:N scheduler) Grows/shrinks dynamically
Rust std::thread Yes Fixed (configurable)
Rust async task No (executor-managed) None

Rust gets Go-like scalability from async, not from growable stacks. An async fn compiles into a state machine that stores only the locals which must survive across .await points — there is no dedicated stack at all, and an executor like Tokio polls only the tasks that can actually make progress. That’s how a small pool of OS threads runs millions of concurrent tasks.

Worth knowing: pre-1.0 Rust did have green threads with segmented, growable stacks, much like early Go. It was removed before 1.0 — it forced a runtime on every program, hurt C interoperability, and clashed with the zero-cost-abstraction goal. The ecosystem settled on stackless futures plus a user-chosen executor instead.