Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Introduction

Cpp2Rust translates C++ to fully safe Rust automatically. It is a syntax-driven translator based on clang’s AST.

Cpp2Rust’s algorithm is described in the paper Cpp2Rust: Automatic Translation of C++ to Safe Rust published at PLDI 2026.

Overview

Cpp2Rust first parses the input C++ file(s) with clang and produces an AST. It then traverses the AST and emits Rust code as strings, inserting calls to the libcc2rs runtime library where needed (e.g., for raw pointer semantics). Finally, the Rust code is pretty-printed using rustfmt to a single .rs file.

By default the reference counting model is used, which produces fully safe Rust. A generator of unsafe Rust is also available through the --model=unsafe command line argument for debugging and performance comparisons.

Runtime library (libcc2rs)

The generated code relies on a runtime library designed to simplify the translation process. C pointers are converted into the Ptr<T> type provided by libcc2rs. Ptr<T> models C pointer semantics, including null, arithmetic, and aliasing, while satisfying Rust’s borrow checker through checked run-time operations.

Building

Requirements

On Ubuntu, install the required dependencies with:

sudo apt install libclang-22-dev clang++-22 ninja-build cmake
pip install ruff==0.15.22

Build

mkdir build
cd build
cmake -GNinja ..
ninja
ninja check

Usage

Translate a single file

./build/cpp2rust/cpp2rust --file=<file>.cpp -o=<file>.rs

By default, the reference counting model is used (fully safe output). To generate unsafe Rust instead:

./build/cpp2rust/cpp2rust --file=<file>.cpp -o=<file>.rs --model=unsafe

Minimal example. Given hello.cpp:

#include <cstdio>
int main() {
  printf("hello world\n");
  return 0;
}

Running ./build/cpp2rust/cpp2rust --file=hello.cpp -o=hello.rs produces:

pub fn main() {
    std::process::exit(main_0());
}
fn main_0() -> i32 {
    println!("hello world");
    return 0;
}

Compile and run with:

rustc hello.rs -L build/libcc2rs-target/release
./hello

Translate a whole program

First generate a compile_commands.json for your project. With CMake this is one extra flag:

cmake -DCMAKE_EXPORT_COMPILE_COMMANDS=ON ..

Then run:

./build/cpp2rust/cpp2rust --dir=<dir> -o <output>.rs

<dir> must be the directory that contains compile_commands.json.

Test Suite

# Run all tests
ninja check

# Run only the unit tests
ninja check-unit

# Run libcc2rs unit tests
ninja check-libcc2rs

# Run libcc2rs-macros unit tests
ninja check-libcc2rs-macros

# Regenerate expected output for unit tests after intentional changes
REPLACE_EXPECTED=1 ninja check-unit

Overview

Translation rules describe how C++ library APIs are mapped to Rust. Each rule module lives in the rules/ directory and pairs a C++ source file (src.cpp) with its Rust translation for each model (tgt_unsafe.rs and tgt_refcount.rs).

Every rule is expressed as ordinary, compilable C++ and Rust source code, free of anything platform dependent. Both sides are run through real compilers at build time, so a rule that does not compile fails the build, and the platform-specific spellings (for example bool canonicalizing to _Bool) are derived by the compiler on the host rather than written by hand. The same rule sources work on every platform cpp2rust builds on.

Rules go through a build-time compilation pipeline before cpp2rust can use them:

  1. You author a rule module: C++ patterns in src.cpp and Rust targets in tgt_unsafe.rs / tgt_refcount.rs.
  2. At build time, two preprocessors compile the module into Rules IR under <build>/rules/<module>/: cpp-rule-preprocessor compiles the C++ side into ir_src.json, and rule-preprocessor compiles the Rust side into ir_unsafe.json and ir_refcount.json.
  3. At startup, cpp2rust loads the Rules IR files and indexes the rules by the canonical signature of the C++ construct they match.

The rest of this part covers each stage:

  • Rule Format: the files that make up a rule module and how the two models are layered.
  • Writing Rules: how to write rules for functions, methods, operators, types, constants, and variadics.
  • Compat Shims: how macro-based libc APIs like errno and FD_SET are rewritten into matchable function calls.
  • Conventions: naming and style conventions rule authors must follow.
  • The Rule Preprocessors: the two build-time tools that compile rules to the Rules IR.
  • The Rules IR: the JSON format the preprocessors emit.
  • Loading and Matching: how cpp2rust loads the Rules IR and matches rules against the input AST.
  • The Matching Engine: how a candidate rule’s signature is unified against the input.
  • Rule Rewriting: how rule bodies are adapted at application time, in particular the with_mut rewrite.

Rule Format

A rule module is a directory under rules/, usually named after the header or library it covers (rules/unistd, rules/vector, rules/string, …). It contains:

  • src.cpp and/or src.c: the C++ (or C) side of each rule.
  • tgt_unsafe.rs: the Rust targets for the unsafe model.
  • tgt_refcount.rs: the Rust targets for the reference counting model (optional, see below).

A rule is a pair of same-named functions on the two sides. Names determine the rule kind:

  • f1, f2, … are expression rules: they map a C++ call, member access, constructor, or constant to a Rust expression.
  • t1, t2, … are type rules: they map a C++ type to a Rust type.

Expression rules

On the C++ side, an fN function must have a body that is exactly one return statement. The returned expression is the pattern: the preprocessor resolves the callee of that expression (the function, method, constructor, enum constant, or macro being used) and that becomes the rule’s matching key. The function parameters stand for the arguments at the call site.

// rules/unistd/src.cpp
int f4(const char *pathname) { return unlink(pathname); }

On the Rust side, the same-named function gives the replacement expression. Parameters must be named a0, a1, … and correspond positionally to the C++ parameters:

#![allow(unused)]
fn main() {
// rules/unistd/tgt_unsafe.rs
unsafe fn f4(a0: *const libc::c_char) -> i32 {
    libc::unlink(a0)
}
}
#![allow(unused)]
fn main() {
// rules/unistd/tgt_refcount.rs
fn f4(a0: Ptr<u8>) -> i32 {
    match nix::unistd::unlink(a0.to_rust_string().as_str()) {
        Ok(()) => 0,
        Err(__e) => {
            libcc2rs::cpp2rust_errno().write(__e as i32);
            -1
        }
    }
}
}

When the converter encounters unlink(x) in the input, it emits the rule body with the translated x substituted for a0.

Type rules

On the C++ side, a tN rule is a type alias (using or typedef). On the Rust side, it is a zero-argument function whose return type is the mapped Rust type and whose body is the default initializer for that type:

// rules/vector/src.cpp
template <typename T1> using t1 = std::vector<T1>;
#![allow(unused)]
fn main() {
// rules/vector/tgt_unsafe.rs
fn t1<T1>() -> Vec<T1> {
    Vec::new()
}
}

Model layering

The loader always reads ir_unsafe.json first. When translating with the reference counting model, it then overlays ir_refcount.json on top: entries with the same rule name replace the unsafe ones.

This means tgt_refcount.rs only needs to contain the rules that differ from the unsafe model. For example, the __builtin_mul_overflow rule in rules/builtin has a pointer out-parameter (a2 below), so the two models translate it differently: in the unsafe model a2 is a raw *mut i64 written through a deref, while in the refcount model it is a Ptr<i64> written through Ptr::write. The other two arguments are identical in both models:

#![allow(unused)]
fn main() {
// rules/builtin/tgt_unsafe.rs
unsafe fn f9(a0: i64, a1: i64, a2: *mut i64) -> bool {
    let (val, ovf) = a0.overflowing_mul(a1);
    *a2 = val;
    ovf
}
}
#![allow(unused)]
fn main() {
// rules/builtin/tgt_refcount.rs
fn f9(a0: i64, a1: i64, a2: Ptr<i64>) -> bool {
    let (val, ovf) = a0.overflowing_mul(a1);
    a2.write(val);
    ovf
}
}

The module’s other rules (byte swaps, __builtin_expect, …) translate identically in both models, so they appear only in tgt_unsafe.rs and the refcount model inherits them. A module where no rule needs a refcount-specific translation can omit tgt_refcount.rs entirely.

C and C++ sources

A module may have both src.c and src.cpp; both are preprocessed and merged into one ir_src.json. Defining the same rule name in both files is a hard error, so numbering must not collide.

This split is necessary because rules match on the exact canonical signature of the callee, and some libc functions have different signatures in C and C++. For example, C has a single char *strchr(const char *, int), while C++ replaces it with const-correct overloads such as const char *strchr(const char *, int). Since the signatures differ, rules/cstring defines one rule per language:

// rules/cstring/src.cpp
const char *f6(const char *a0, int a1) { return strchr(a0, a1); }
// rules/cstring/src.c
char *f5(const char *a0, int a1) { return (strchr)(a0, a1); }

The C++ rule matches strchr calls in code translated as C++, the C rule matches them in code translated as C.

rules/builtin uses this to cover both languages: src.cpp defines f9/f10 for the C++ __builtin_mul_overflow (returning bool) while src.c defines f12/f13 for the C version (returning int); their Rust bodies are identical.

The rules crate

The whole rules/ tree is a single Rust crate. rules/build.rs walks the tree, collects every tgt_*.rs, and generates rules/src/modules.rs with one #[path = ...] module per file. Building the crate therefore type-checks every rule body against the crates the rule targets call into, which are declared as dependencies in rules/Cargo.toml (libcc2rs, libc, nix, …). The Rust rule preprocessor compiles exactly this crate to resolve types in rule bodies. rules/src/ is the only subdirectory that is not a rule module.

Writing Rules

This page shows how to write rules for each kind of C++ construct. In every case the recipe is the same: write an fN (or tN) function on the C++ side whose single return statement exercises the construct, and a same-named function on the Rust side giving the translation.

Free functions

// rules/stat/src.cpp
int f1(const char *pathname, struct stat *statbuf) {
  return stat(pathname, statbuf);
}
#![allow(unused)]
fn main() {
// rules/stat/tgt_refcount.rs
fn f1(a0: Ptr<u8>, a1: Ptr<Stat>) -> i32 {
    match nix::sys::stat::stat(a0.to_rust_string().as_str()) {
        Ok(__s) => {
            a1.with_mut(|__st| *__st = Stat::from_libc(&__s));
            0
        }
        Err(__e) => {
            libcc2rs::cpp2rust_errno().write(__e as i32);
            -1
        }
    }
}
}

Rule bodies may be arbitrarily complex; multi-statement bodies are wrapped in a block when spliced into the output.

return statements are prohibited in Rust rule bodies (the preprocessor rejects them); produce the result as a tail expression instead. The body is not emitted as a function of its own: it is spliced inline into the generated code as a block expression, so a return would not end the rule, it would return from whatever generated function the rule happens to be expanded in.

When the pattern’s type cannot be named, the rule uses an auto return type: rules/iomanip writes auto f1(int n) { return std::setw(n); } because std::setw returns an unspecified type.

Methods

There is no special syntax for member functions: write a free function that takes the receiver as its first parameter and calls the method on it. On the Rust side the receiver is a0.

// rules/vector/src.cpp
template <typename T1> std::size_t f2(const std::vector<T1> &o) {
  return o.size();
}
#![allow(unused)]
fn main() {
// rules/vector/tgt_unsafe.rs
unsafe fn f2<T1>(a0: Vec<T1>) -> usize {
    a0.len()
}
}

Template rules use generic parameters named T1, T2, … on both sides, matched positionally. The rule is written against the open template std::vector<T1>, with T1 left as a placeholder, so a single rule covers every instantiation: when the input program calls size() on, say, a std::vector<int>, the matcher binds T1 = int.

Static member functions

A static member function is also written with a receiver parameter, which exists only to name the class. The call site has no receiver argument, so the Rust side drops it and numbers the remaining parameters from a0; here there are none:

// rules/limits/src.cpp
template <typename T1> T1 f1(std::numeric_limits<T1> &a0) { return a0.max(); }
#![allow(unused)]
fn main() {
// rules/limits/tgt_unsafe.rs
unsafe fn f1<T1: HasMinMax>() -> T1 {
    <T1>::MAX
}
}

(HasMinMax is a helper trait defined alongside the rules in the same file.)

Constructors

Constructors are functions returning the type by value, one rule per overload:

// rules/string/src.cpp
std::string f7(const char *s, std::size_t n) { return std::string(s, n); }
std::string f9(std::size_t n, char ch) { return std::string(n, ch); }

Overloads that differ in value category are distinct rules too: rules/vector has separate rules for push_back(const T1 &) and push_back(T1 &&).

No destructor rules exist so far: the STL and libc APIs covered by the current rules have not needed any, since their types map to Rust types whose Drop implementations already do the right thing.

Operators

Write operators with explicit operator call syntax, in member form (x.operator@(...)) or free form (operator@(a, b)):

// rules/map/src.cpp
template <typename T1, typename T2>
T2 &f1(std::map<T1, T2> &o, const T1 &key) { return o.operator[](key); }

template <typename T1, typename T2>
bool f11(typename std::map<T1, T2>::iterator a,
         typename std::map<T1, T2>::iterator b) {
  return operator!=(a, b);
}

Post-increment is distinguished from pre-increment by the usual dummy int parameter: a0.operator++(a1) versus it.operator++(). Conversion operators use the same explicit syntax: a0.operator T1 &() in rules/functional matches the conversion of a std::reference_wrapper<T1> back to a reference. Field accesses are rules of their own, matched by the field: it->first and it->second through iterators, plain o.second on a pair (rules/map, rules/pair).

Callable arguments

A rule parameter may be a callable. Function pointers are spelled directly; for a lambda, whose type cannot be written, the rule declares a file-scope lambda and takes decltype(lambda):

// rules/algorithm/src.cpp
auto lambda = [](const T2 &a, const T2 &b) { return false; };
void f6(T1 first, T1 last, decltype(lambda) comp) {
  return std::stable_sort(first, last, comp);
}
#![allow(unused)]
fn main() {
// rules/algorithm/tgt_unsafe.rs
unsafe fn f6<T1: Ord, T2>(a0: *mut T1, a1: *mut T1, a2: &mut T2)
where
    T2: FnMut(&T1, &T1) -> bool,
{ ... }
}

T1 and T2 are not template parameters here but file-scope helper structs modelling an iterator and its value type; being named like generics, they bind as T1/T2 at the use site. The function pointer version of the comparator is a separate rule (f7).

Iterators

There is no iterator abstraction: an iterator type gets a type rule, and every operation on it its own expression rule (operator*, operator++, operator!=, …). What the type maps to is up to the rule: std::string::iterator becomes a plain pointer (*mut libc::c_char unsafe, Ptr<u8> refcount), while std::map iterators become the runtime types libcc2rs::UnsafeMapIterator/MapIterator. Dependent iterator types are named with typename:

// rules/map/src.cpp
template <typename T1, typename T2>
using t2 = typename std::map<T1, T2>::const_iterator;

Types

A type rule has two halves. On the C++ side, declare a type alias named tN for the C++ type being mapped. On the Rust side, write a function with the same name that takes no arguments: its return type is the Rust type that the C++ type maps to, and its body is the default value the generated code uses when it needs to construct one (e.g. for an uninitialized variable). Reference and pointer variants of a type each get their own rule:

// rules/iostream/src.cpp
using t1 = std::ostream;
using t2 = std::ostream &;
using t3 = std::ostream *;

C structs use typedef instead of using:

// rules/stat/src.cpp
typedef struct stat t1;
#![allow(unused)]
fn main() {
// rules/stat/tgt_unsafe.rs
fn t1() -> ::libc::stat { unsafe { std::mem::zeroed() } }
}
#![allow(unused)]
fn main() {
// rules/stat/tgt_refcount.rs
fn t1() -> libcc2rs::Stat { Default::default() }
}

A type rule may map to the sentinel type libcc2rs::IgnoreRule, meaning “this model has no special mapping for the type”; the converter then falls back to its normal type conversion. This is useful when only one model needs a custom mapping: rules/carray maps multi-dimensional C arrays to nested boxed slices in the refcount model, while its tgt_unsafe.rs targets are IgnoreRule so the unsafe model keeps the default array conversion.

Enum values, constants, and macros

Constants are fN functions that take no arguments and return the constant, one rule per value:

// rules/fcntl/src.cpp
int f3(void) { return O_CREAT; }
int f4(void) { return O_TRUNC; }
#![allow(unused)]
fn main() {
// rules/fcntl/tgt_unsafe.rs
unsafe fn f3() -> i32 { ::libc::O_CREAT }
}

For macros that expand to integer literals, the preprocessor records the macro name rather than the value, so O_CREAT in the input matches this rule by name. Enum constants and global variables (e.g. std::cout) are matched by their qualified name. A global and its address are separate rules: rules/iostream maps both std::cout (f1) and &std::cout (f3).

A variable of class type, such as std::cout or std::strong_ordering::less, must be returned by reference. Returned by value, the pattern is a copy construction, so the rule keys on the copy constructor instead of the variable and matches every copy of that type:

// rules/compare/src.cpp
const std::strong_ordering &f1() { return std::strong_ordering::less; }
const std::strong_ordering &f2() { return std::strong_ordering::equal; }

The Rust side still returns the value; the reference only exists to keep the C++ pattern free of the copy.

Integer-literal macros are the only macros matchable directly. Macros whose expansions are platform internals with no stable callee, such as errno or FD_SET, are first rewritten into calls to synthetic cpp2rust_* functions by the compat shims; rules then match the shim call.

Variadic functions

The C++ side uses a template parameter pack rather than a C-style ... parameter, out of necessity: a function that takes ... cannot forward its variadic arguments to another call, so a rule like

int f1(int a0, int a1, ...) { return fcntl(a0, a1, ...); }

is not expressible. A parameter pack can be forwarded (args...), which is exactly what the rule body needs to do. The Rust side takes a trailing parameter that must be typed &[VaArg] and named va:

// rules/fcntl/src.cpp
template <typename... Args>
int f1(int a0, int a1, Args... args) {
  return fcntl(a0, a1, args...);
}
#![allow(unused)]
fn main() {
// rules/fcntl/tgt_refcount.rs
fn f1(a0: i32, a1: i32, va: &[VaArg]) -> i32 { ... }
}

Bodies read the arguments through the va-args API in libcc2rs (VaArg, VaList, the VaArgGet accessors, format_c).

Constructing from forwarded arguments

Functions like emplace_back forward their arguments to a constructor. Rust has no equivalent, so the rule receives the finished value instead: the Rust side takes a trailing parameter named init, and the converter builds it at the call site from the arguments after the fixed ones. The C++ side spells the pack as Init<T, Args>, a transparent alias (template <typename T, typename A> using Init = A;) whose T names the type to build:

// rules/vector/src.cpp
template <typename T1, typename... Args>
T1 &f112(std::vector<T1> &o, Init<T1, Args> &&...args) {
  return o.emplace_back(std::forward<Args>(args)...);
}
#![allow(unused)]
fn main() {
// rules/vector/tgt_unsafe.rs
unsafe fn f112<T1>(a0: &mut Vec<T1>, init: T1) {
    let __init = init;
    a0.push(__init)
}
}

T must be one of the callee’s template arguments. The preprocessor records its position as a (depth, index) pair, and at a call like v.emplace_back(4, 5) on a std::vector<Point> the converter reads the template argument at that position from the resolved callee (Point). It then asks Sema which constructor builds a Point from (4, 5) and substitutes the converted construction for init. Binding init to a local before touching a0 keeps the construction from overlapping a borrow of the container.

Passthrough rules

When a call should be forwarded verbatim to the same-named function in Rust’s libc crate, the Rust target can be an extern declaration instead of a body:

// rules/fcntl/src.cpp
template <typename... Args>
int f1(int a0, int a1, Args... args) {
  return fcntl(a0, a1, args...);
}
#![allow(unused)]
fn main() {
// rules/fcntl/tgt_unsafe.rs
unsafe extern "C" {
    fn f1(a0: i32, a1: i32, ...) -> i32;
}
}

The converter then emits a direct libc::fcntl(...) call at the call site.

Platform-specific rules

Gate the C++ side with the usual preprocessor conditionals and the Rust side with #[cfg(...)]; the two must agree so that the rule name sets line up:

// rules/socket/src.c
#ifdef __linux__
int f4(void) { return SOCK_CLOEXEC; }
#endif
#![allow(unused)]
fn main() {
// rules/socket/tgt_unsafe.rs
#[cfg(target_os = "linux")]
unsafe fn f4() -> i32 {
    libc::SOCK_CLOEXEC
}
}

The Rust preprocessor evaluates #[cfg] attributes against the host target (only target_os = linux|macos and target_arch = x86_64|x86 are accepted) and drops non-matching rules.

Mutually exclusive platform branches use #elif with disjoint rule numbers: rules/errno defines f91 to f135 under __linux__ and f136 to f153 under __APPLE__. Feature-test macros a pattern needs must come before the includes, as with #define _GNU_SOURCE in rules/socket/src.c.

Pattern resolution limits

The preprocessor resolves a template pattern by instantiating its template parameters with synthesized types (The Rule Preprocessors):

  • A bare T1 becomes an empty struct, so the pattern cannot use members, operators, or nested types of T1.
  • A parameter pack instantiates to the empty pack.
  • A non-type parameter is pinned to the value 1.

Unqualified callees are looked up in namespace std first and in the global scope only when std has no match, so an unqualified name that exists in both resolves to the std one.

Compat Shims

Rule matching needs a resolvable callee. The preprocessor keys every expression rule on the function, method, constructor, constant, or global that the pattern’s return expression resolves to, and the only macros it can record are those that expand to an integer literal, which match by macro name. Any other macro is invisible to the rule system: by the time clang has built the AST, the macro is gone and only its expansion remains.

That is a problem for a small set of libc APIs that are specified as macros over platform internals:

  • errno is an object-like macro; glibc expands it to (*__errno_location()), macOS to (*__error()).
  • assert expands to a conditional that stringifies the condition and calls a platform-specific failure handler with file and line arguments.
  • FD_SET, FD_CLR, FD_ISSET, and FD_ZERO expand to bit manipulation on the fd_set representation, through helpers that differ per platform.
  • ntohl, ntohs, htonl, and htons expand to byte-swap builtins or to nothing at all, depending on endianness.

There is no stable, platform-independent callee here to key a rule on. The compat headers in cpp2rust/compat/ fix this by rewriting each such macro into a call to a synthetic, well-known function before matching happens.

How the shims work

cpp2rust/compat is injected as a system include directory ahead of the platform headers in every clang invocation the project makes: both when cpp2rust parses the input program and when cpp-rule-preprocessor compiles rule sources. The shared flag list lives in cpp2rust/compat/platform_flags.h (getPlatformClangBeginFlags), and the directory path is baked in at build time via the COMPAT_INCLUDE_DIR definition.

A shim header sits at the same relative path as the real header it shadows (errno.h, sys/select.h, arpa/inet.h, …), so an ordinary #include <errno.h> finds the shim first. The header then:

  1. pulls in the real platform header with #include_next (a GNU extension; the shared flags pass -Wno-gnu-include-next for it),
  2. #undefs the macro,
  3. declares a cpp2rust_* shim function,
  4. redefines the macro to call the shim.

cpp2rust/compat/errno.h in full:

#include_next <errno.h>

#undef errno

int *cpp2rust_errno(void);

#define errno (*cpp2rust_errno())

The redefinition keeps errno an lvalue by dereferencing the returned pointer, so both reads and assignments like errno = 0 still parse; what the matcher sees in either case is a call to int *cpp2rust_errno().

Because the input program and the rule sources are compiled with the same shim headers, both sides canonicalize to the same signature, and an ordinary expression rule matches it:

// rules/errno/src.c
#include <errno.h>

int *f1(void) { return cpp2rust_errno(); }

The shim functions are declared but never defined on the C side. They only exist so that the callee resolves; translation replaces the call with the rule body, so no C implementation is ever linked. Whatever the shim is supposed to do is supplied by the Rust targets:

#![allow(unused)]
fn main() {
// rules/errno/tgt_unsafe.rs
unsafe fn f1() -> *mut i32 {
    libcc2rs::cpp2rust_errno_unsafe()
}
}
#![allow(unused)]
fn main() {
// rules/errno/tgt_refcount.rs
fn f1() -> Ptr<i32> {
    libcc2rs::cpp2rust_errno()
}
}

In the unsafe model libcc2rs::cpp2rust_errno_unsafe wraps the real platform errno location (__errno_location on Linux, __error on macOS). The refcount model instead virtualizes errno as a thread-local Value<i32> inside libcc2rs. Nothing else writes that cell, so it is a discipline of the rules: every refcount rule that translates a call that can fail must write the error code into it on the failure path, as libcc2rs::cpp2rust_errno().write(__e as i32) in the stat rule; a rule that skips the write breaks programs that check errno.

A rule pattern may spell either the macro or the shim directly; the two are identical after expansion. rules/errno and rules/assert call the shim by name, while rules/arpa_inet and rules/select write the macro form:

// rules/select/src.cpp
void f2(int fd, fd_set *set) { return FD_SET(fd, set); }

The current shims

HeaderMacrosShim functionsRules
assert.hassertcpp2rust_assert_fail(bool)rules/assert maps it to assert!(a0)
errno.herrnocpp2rust_errno()rules/errno, see above
arpa/inet.hntohl, ntohs, htonl, htonscpp2rust_ntohl(x), …rules/arpa_inet maps them to u32::from_be, u16::to_be, …
sys/select.hFD_SET, FD_CLR, FD_ISSET, FD_ZEROcpp2rust_fd_set(fd, set), …rules/select maps them to libc::FD_SET(...) (unsafe) or CFdSet methods (refcount)

Note how the shim also normalizes the shape of the API. C’s assert is a macro precisely so it can stringify its condition and capture file and line; the shim reduces it to a plain void(bool) function, and the Rust side regains the diagnostics by mapping it to the assert! macro.

Adding a new shim

To make another macro-based API matchable:

  1. Create the header in cpp2rust/compat/ at the same relative path as the platform header that defines the macro.
  2. Follow the pattern above: #include_next the real header, #undef the macro, declare a cpp2rust_<name> function with the macro’s effective signature, and redefine the macro to call it.
  3. Write rules for the shim in a rules/ module as for any other function, including the corresponding header in src.c/src.cpp.
  4. If a model needs runtime support (as refcount errno does), implement it in libcc2rs and call it from the rule target.

Keep the shim’s signature platform-independent; the whole point is that both sides of every rule see one canonical declaration on every platform.

The same shared flag list also passes -D_FORTIFY_SOURCE=0, which keeps glibc from substituting fortified variants (__printf_chk and friends) for standard calls. Like the shims, this ensures that calls in the input program resolve to the standard declarations the rules are written against.

Conventions

Most of these conventions are enforced by the preprocessors, and violating them fails the build; the notes below call out the ones that are not checked.

Naming

ElementC++ sideRust side
Expression rulef1, f2, …same name
Type rulet1, t2, … via using/typedeffn tN() -> RustType with no arguments
Parametersfree-form (o, it, key, dst, n, …)must be a0, a1, … consecutive from 0
GenericsT1, T2, … (type and non-type params)T1, T2, … consecutive from 1
Variadic packtypename... Argstrailing va: &[VaArg]
Constructing packInit<T, Args> &&...argstrailing init: RustType
Locals in Rust bodiesdouble-underscore prefix: __v, __fd, __e, …

Notes:

  • Rule numbering is per module, and gaps are currently allowed (e.g. rules/map has no f4), though this might change in the future. Names must be unique across src.c and src.cpp combined.
  • On the C++ side parameter names are free, but the order defines the placeholder indices: the first parameter is a0 on the Rust side, the second is a1, and so on. The receiver of a method rule is always the first parameter, hence a0.
  • Generic parameters are matched positionally between the two sides, so T1 in the Rust target means “whatever bound to T1 in the C++ pattern”.
  • Locals introduced inside Rust rule bodies use a __ prefix. This is not checked by the build, but it is needed: rule bodies are spliced inline into the generated code, so an unprefixed local could collide with a variable name from the translated program.

Function qualifiers

  • In tgt_unsafe.rs, expression rules are unsafe fn; type rules (tN) are plain fn.
  • In tgt_refcount.rs, all rules are safe fn. The refcount model produces fully safe Rust, so a refcount rule body must not need unsafe.

The build does not check the qualifiers themselves; only rustc’s usual rules apply when the rules crate compiles. In particular, nothing stops an unsafe block inside a refcount rule body from being spliced into the output, so keeping refcount rules safe is what upholds the model’s safety guarantee.

C++ pattern shape

  • An fN body must be exactly one return statement. The preprocessor rejects anything else.
  • return statements are not allowed inside Rust rule bodies; write the result as a tail expression instead.
  • Exercise exactly one construct per rule. If an API has several overloads, write one rule per overload (including separate rules for const T & versus T && parameters).

Argument accesses

Every use of an aN parameter in a rule body is classified as a read, write, or move by the rule preprocessor. Passing an argument by value counts as a read, not a move; the only way to record a move is std::mem::take(&mut aN).

Type checking

All tgt_*.rs files are compiled as part of the rules crate, so a rule body that does not type-check against libcc2rs, libc, nix, etc. breaks the build. If a rule needs a new crate dependency, add it to rules/Cargo.toml and to the hardcoded crate list in rule-preprocessor/src/semantic.rs (see The Rule Preprocessors).

The Rule Preprocessors

Two build-time tools compile rule modules into the Rules IR that cpp2rust loads at runtime. Both write into <build>/rules/<module>/:

  • cpp-rule-preprocessor compiles src.cpp/src.c into ir_src.json.
  • rule-preprocessor compiles tgt_unsafe.rs/tgt_refcount.rs into ir_unsafe.json/ir_refcount.json.

The C++ side is keyed by resolved callee signatures; the Rust side by rule names. The two are joined by rule name when cpp2rust loads them.

cpp-rule-preprocessor

A clang LibTooling executable (cpp2rust/cpp_rule_preprocessor.cpp) that runs once per rule directory:

cpp-rule-preprocessor --dir rules/string --out <build>/rules/string/ir_src.json

Extra compiler flags can be passed with repeated --cxxflags options, though CMake, which invokes the tool for every rule module via the preprocess-cpp-rules target, passes none. Note that the parent directory of --out must already exist; CMake creates it before each invocation, so a manual run must do the same.

Rule sources are always compiled with the fixed flag set from cpp2rust/compat/platform_flags.h, the same one used to parse input programs (see Compat Shims). There is no compilation database and no -std= flag: the language is chosen by clang from the file extension, and src.c is processed before src.cpp.

For each rule it:

  1. Validates that every fN body is exactly one return statement.
  2. Resolves the callee of the returned expression. For non-template rules this is just the called declaration. For template rules the callee is unresolved, so the tool instantiates the rule’s template parameters with synthesized dummy types and runs overload resolution to find the function the rule refers to.
  3. Prints the resolved declaration as a canonical signature string: <return type> <qualified::name>(<param types>[, ...])[ const][ volatile][ &|&&], where , ... appears for C-variadic functions and the trailing qualifiers only for methods. For tN aliases it prints the underlying type.

The output is a flat JSON object mapping rule names to these signature strings.

The printer preserves typedef sugar instead of desugaring it: size_t prints as size_t, not unsigned long, which is what lets it map to usize while plain unsigned long maps to u64 (for tN aliases this preservation is explicit; inside function signatures the spelling survives through the printing policy). Integer literals expanded from a macro are recorded as the macro name, which is how constant rules like the O_CREAT one match by name.

rule-preprocessor

A Rust binary crate built with the nightly toolchain because it links the compiler’s own libraries (rustc_driver, rustc_middle, …). It processes the whole rules tree in one invocation:

CARGO_TARGET_DIR=<target> cargo +nightly build --release \
    --message-format=json-render-diagnostics \
    --manifest-path rule-preprocessor/Cargo.toml > <target>/artifacts.json
RULE_PREPROCESSOR_ARTIFACTS=<target>/artifacts.json \
CARGO_TARGET_DIR=<target> cargo +nightly run --release \
    --manifest-path rule-preprocessor/Cargo.toml -- <build>/rules [rules-dir]

The environment is load-bearing:

  • RULE_PREPROCESSOR_ARTIFACTS must be set (the tool aborts otherwise): it points to cargo’s JSON build output, from which the rlibs of the rule dependencies (libcc2rs, libc, nix, …) are taken. The target dir is not scanned directly because it may hold stale copies of the same crates with different hashes (e.g., from cargo clippy or an older toolchain), and mixing them makes type checking fail. The crate list is hardcoded, so a new dependency in rules/Cargo.toml also needs an entry in rule-preprocessor/src/semantic.rs.

  • The sysroot comes from running rustc --print=sysroot, so the rustc on PATH must be the same nightly the preprocessor was built with (running through cargo +nightly run guarantees this).

  • rules-dir is optional and defaults to the relative path ../rules, resolved against the current working directory of the process.

CMake drives all of this via the preprocess-rust-rules target: it first builds the rules crate with the stable toolchain (which also regenerates rules/src/modules.rs), then builds the preprocessor in <build>/rule-preprocessor-target, saving cargo’s output to artifacts.json there, and runs it. That initial cargo build of the rules crate is what actually gates the build on rule bodies type-checking (see below). The preprocessor works in two phases.

Phase 1, syntactic. Each tgt_*.rs file is parsed with rust-analyzer’s parser, and functions whose #[cfg] does not match the host are dropped. Every function body is then turned into a list of fragments, whose kinds are described in The Rules IR. The fragmentation is mainly concerned with how the rule’s arguments are used: references to parameters and generics become placeholder and generic fragments, while source text that does not involve an argument is kept as-is.

Each placeholder is tagged with an access: read, write, or move. Some uses give the access away syntactically (&mut a0 is a write); those that do not, typically method-call receivers and arguments, are left as unknown for phase 2. This phase also applies the two preprocessor-side rewrites that support rule rewriting.

Phase 2, semantic. The preprocessor compiles the rules crate in-process with rustc and walks the typed HIR. This gives it the real signature of every callee, which resolves the unknown accesses: passing to a &mut/*mut parameter is a write, to a &/*const parameter a read, and to std::mem::take a move. For type rules it also records which of the nine derivable standard traits (Copy, Clone, Debug, Default, PartialEq, Eq, PartialOrd, Ord, Hash) the mapped type implements. A placeholder still unknown after this phase fails the build.

The preprocessor assumes the rules crate is buildable, which the earlier cargo build of the crate ensures; errors from the in-process compilation are therefore only reported as a warning.

The result is one ir_<model>.json per input file, keyed by rule name. The output file name is derived from the input file name (tgt_unsafe.rs becomes ir_unsafe.json) and the module directory is the direct parent of the tgt_*.rs file.

The Rules IR

Each rule module compiles to up to three JSON files in <build>/rules/<module>/:

  • ir_src.json: the C++ side, from cpp-rule-preprocessor.
  • ir_unsafe.json: the Rust side for the unsafe model, from rule-preprocessor.
  • ir_refcount.json: the Rust side for the refcount model, also from rule-preprocessor (only if the module has a tgt_refcount.rs).

All three are objects keyed by rule name (f1, t1, …), and the loader joins them by name.

Source IR (ir_src.json)

A flat map from rule name to the canonical signature of the C++ construct the rule matches. For rules/vector:

{
  "t1": "std::vector<T1>",
  "f3": "_Bool std::vector<T1>::empty() const"
}

This signature string is the lookup key for the whole rule: the converter prints C++ constructs from the input AST with the same printer and compares the strings.

Parameter packs

A function parameter pack prints as &&..., so the key of a rule for emplace_back(Args &&...args) matches calls with any number of arguments. A rule that takes an init value also records which template argument of the callee init builds:

"f112": {
  "key": "T1 & std::vector<T1>::emplace_back(&&...)",
  "init_type": { "depth": 0, "index": 0 }
}

Target IR (ir_unsafe.json / ir_refcount.json)

An expression rule serializes as an ExprRule object: the rule’s signature plus its body as a list of fragments. For unsafe fn f6<T1>(a0: &mut Vec<T1>) -> *mut T1 { a0.as_mut_ptr() }:

"f6": {
  "body": [
    { "method_call": {
        "receiver": [ { "placeholder": { "arg": 0, "access": "read" } } ],
        "body": [ { "text": ".as_mut_ptr()" } ] } }
  ],
  "generics": { "T1": [] },
  "params": { "a0": { "type": "&mut Vec<T1>" } },
  "return_type": { "type": "*mut T1", "is_unsafe_pointer": true }
}

The fragment kinds are:

  • text: literal Rust source, emitted verbatim.
  • placeholder: a use of one of the rule’s aN parameters in the body (not an argument of whatever the body calls); the converter substitutes the translated call-site argument here. Its fields:
    • arg: the parameter index N.
    • access: how the body uses the argument: read, write, or move.
    • is_index_base: the placeholder is the base of an index expression.
  • generic: a TN slot, replaced with the instantiated Rust type; serialized as the 1-based index N.
  • method_call: a method call split into receiver and body fragment lists, so the code generator can rewrite the pair (see Rule Rewriting).
  • va_args: the expansion point for a variadic tail.
  • init: the value built from the call’s trailing arguments.

Every type in the Rules IR (in params, return_type, and type rules) is a TypeInfo object, the type text plus a set of flags:

  • is_refcount_pointer: the type is a Ptr<...>.
  • is_unsafe_pointer: the type is a raw *mut/*const pointer.
  • derives (type rules only): the standard traits the mapped type implements (Copy, Clone, Default, …).

The two pointer flags are mutually exclusive; the loader rejects a type with both set.

An ExprRule carries two flags of its own:

  • multi_statement: the body has more than one statement, or a statement followed by a tail expression, and must be wrapped in a block to stay a single expression.
  • is_extern: the rule is an extern passthrough declaration and has no body.

Fields that are false, empty, or unset are omitted from the Rules IR. A va or init parameter is never listed in params, and a () return type is omitted.

A type rule serializes as a TypeRule object: its TypeInfo plus the init initializer expression, merged into one object:

"t1": { "type": "Vec<T1>", "init": "Default::default()" }

There is no explicit tag distinguishing the two rule kinds: an entry with body is an expression rule, one with type and init a type rule.

In-memory form

cpp2rust mirrors the Rules IR in C++ structs of the same names, defined in cpp2rust/converter/translation_rule.h. TranslationRule::Load reads one module directory (the ir_*.json files described above) and produces two maps keyed by rule name, one holding ExprRules and one holding TypeRules:

  • An ExprRule holds the body fragments, the parameter and return TypeInfos, and the two rule-level flags (multi_statement and is_extern). The name-keyed Rules IR maps become positional vectors: parameter aN is entry N of params, generic TN is entry N-1 of generics (each entry being the bound list). Rules support at most 9 generic parameters (kMaxGenerics).
  • A TypeRule holds the mapped type’s TypeInfo and the initializer expression. The same struct also represents the built-in type mappings (scalars, pointers, …) that the loader registers directly in code, without any Rules IR behind them: for example, int maps to i32, and int * to *mut i32 in the unsafe model or Ptr<i32> in the refcount model.

Both also carry src, the canonical C++ signature attached from ir_src.json; it is the key the rule is matched by.

How the loader finds the Rules IR directory, overlays the refcount model on the unsafe one, and indexes the loaded rules for matching is covered in Loading and Matching.

Loading and Matching

Finding the rules directory

cpp2rust takes the Rules IR directory via --rules <dir>. If the flag is omitted it tries ./rules and then <executable dir>/../rules, accepting the first candidate that contains (recursively) a subdirectory with ir_src.json plus ir_unsafe.json or ir_refcount.json. Since the build writes the Rules IR to <build>/rules and the binary lands in <build>/bin, the default resolution picks up the generated Rules IR without any flags.

Loading

Rules are loaded once per process by Mapper::LoadTranslationRules:

  1. Built-in type mappings are registered first. Every scalar is mapped with its width taken from the host (int maps to i32, unsigned long to u64), together with its const form and its pointer forms: *mut/*const in the unsafe model, Ptr<T> in the refcount model, where constness is dropped. char maps to libc::c_char in the unsafe model and to u8 in refcount; size_t/ssize_t map to usize/isize; void * maps to *mut ::libc::c_void, or in refcount to AnyPtr.
  2. Every subdirectory of the rules directory is loaded with TranslationRule::Load, which reads ir_unsafe.json, overlays ir_refcount.json when translating with the refcount model, and then attaches the C++ signature from ir_src.json to each rule by name.

Loading is strict: an ir_src.json entry with no matching target rule is a fatal error (this is what catches mismatched #if/#[cfg] gating), every generic declared by a rule must appear in its C++ signature, and two type rules for the same C++ type are rejected.

Matching

Loaded rules are indexed in two multimaps, one for expression rules and one for type rules. The multimap key is only a coarse bucket for collecting candidate rules; whether a candidate actually matches is decided by the matching engine. The bucket key is derived from the C++ signature:

  • For expressions, the qualified function name with the return type, the parameter list, and all template arguments stripped, so the rule for _Bool std::vector<T1>::empty() const lands in the std::vector::empty bucket.
  • For types, the text up to the first <, so all std::vector<...> rules land in the std::vector bucket.

During translation, the converter prints the construct it encounters with the same canonical printer used by cpp-rule-preprocessor, which is what makes the two sides comparable:

  • Functions and methods print as <return type> <qualified::name>(<param types>[, ...])[ const][ volatile][ &|&&].
  • Enum constants and global variables print as their qualified name.
  • Integer literals expanded from a macro print as the macro name.

All rules in the matching bucket are then unified against this string by the matching engine, which binds T1…T9 to the concrete types at the use site and picks the most specific rule when several match. Type lookups first try the sugar-preserving spelling (so a rule can match size_t as written) and retry with the desugared type on failure.

Running cpp2rust with --verbose logs every lookup and the rule it matched, which is the quickest way to see why a rule does or does not fire.

Application

When a rule matches, the converter walks its body fragments and emits:

  • text fragments verbatim,
  • placeholder fragments as the translated call-site argument. How the argument is emitted depends on the placeholder’s access and on whether the argument and the declared parameter type are pointers:
    • Read access emits the argument as a plain value, with an implicit numeric cast when the parameter type asks for one.
    • Write access emits the argument as an lvalue.
    • Move access wraps the argument in std::mem::take(&mut ...); temporaries are moved as-is.
    • If the rule declares a pointer parameter but the argument is not a pointer, the converter takes a fresh pointer to it, materializing a temporary when the argument has no address of its own. For example, the refcount std::max rule declares Ptr<T1> parameters, since the C++ side takes const T1 & and the refcount model translates references as Ptr, so std::max(x1, x2) on plain locals substitutes x1.as_pointer() and x2.as_pointer() for the placeholders, while std::max(30, 40) first materializes __tmp_0 and __tmp_1 values for the literals and points into those.
    • If the receiver argument is a pointer but the rule expects a value, the converter dereferences it; this only happens for receivers, not for ordinary arguments.
  • generic fragments as the Rust mapping of the bound C++ type,
  • va_args fragments as the converted variadic tail,
  • init fragments as the value built from the call’s trailing arguments,
  • method_call fragments as receiver followed by body, possibly rewritten (see Rule Rewriting).

Multi-statement bodies are wrapped in { } so they remain a single expression. Rules for user-defined C++ types are injected through the same mechanism at translation time.

The Matching Engine

Loading and Matching collects the candidate rules for a construct from a bucket; each candidate’s source signature is then unified against the printed construct. The signature is treated as a template whose T1…T9 slots capture concrete types: std::vector<T1>::vector() unifies with std::vector<int>::vector() by binding T1 = int.

Unification works on the two strings:

  • Whitespace differences are ignored.
  • A TN slot captures up to the next literal text of the pattern, found at the same <>/()/[] nesting depth. This is how T1 captures all of std::map<int, int> in std::vector<std::map<int, int>> without stopping at the inner comma.
  • A TN that appears again must match its first capture exactly.
  • The whole printed string must be consumed; trailing text fails the match.
  • Slots may stay unbound (a pattern can use T2 without T1).

A rule matches if unification succeeds. When several rules in the bucket match, the one with the longest source signature wins, so more specific rules take precedence; between equally long signatures the choice is unspecified.

Bucket keys

The bucket keys described in Loading and Matching have two special cases: array types bucket by the text after the first [ rather than the text before a <, and operator() rules are cut at the operator’s own parentheses, so their key ends at ...::operator.

Instantiating the target

Captures are C++ spellings. Before being substituted into the rule’s Rust fragments, each capture is itself mapped through the type rules, recursively, so T1 = std::vector<int> substitutes as Vec<i32>. A captured type with no type rule of its own is an error.

Rule Rewriting

A rule body is written against idiomatic Rust types: a rule that mutates a vector declares its parameter as &mut Vec<T1>. But in the refcount model the call-site argument is usually a Ptr<Vec<T1>>, and a Ptr cannot produce a long-lived &mut. Instead of forcing every rule to handle pointers, the code generator rewrites the rule body at application time.

The with_mut rewrite

libcc2rs provides

#![allow(unused)]
fn main() {
impl<T> Ptr<T> {
    pub fn with_mut<R>(&self, f: impl FnOnce(&mut T) -> R) -> R { ... }
}
}

which checks the pointer, borrows the pointee mutably, and runs the closure on it (with an immutable sibling Ptr::with). The refcount converter uses it to bridge the gap; the unsafe converter never rewrites and simply emits receiver followed by body. The rewrite fires when all three hold:

  1. The rule body fragment is a method call whose receiver contains a placeholder (the preprocessor splits every method call into receiver and body fragments precisely to enable this). If the receiver contains several placeholders, the first one is used.
  2. The receiver placeholder’s access is write or move, i.e. the method takes &mut self or the rule mutates the parameter. Read access does not need the rewrite, since a read can go through a StrongPtr obtained with Ptr::upgrade, or through a read() copy.
  3. The call-site argument is a pointer, or an expression of reference type (which includes an operator call returning a reference).

The rule’s method call a0.method(...) is then emitted as

#![allow(unused)]
fn main() {
ptr.with_mut(|__v: <rule param type>| __v.method(...))
}

For example, the push_back rule is written as an ordinary &mut method call:

#![allow(unused)]
fn main() {
fn f21<T1: Clone>(a0: &mut Vec<T1>, a1: T1) { ... a0.push(...) }
}

Given the C++ input v.push_back(20); where v is reached through a Ptr<Vec<i32>>, the generated code is:

#![allow(unused)]
fn main() {
v.with_mut(|__v: &mut Vec<i32>| __v.push(20));
}

When the receiver is a plain local value rather than a pointer, condition 3 fails and no closure is emitted; the same rule produces a direct call like (*v2.borrow_mut()).push(0);.

The rewrite applies to pointer dereferences (p->push_back(20)) and to reference usages (r.push_back(20) with std::vector<int> &r = *p); both are translated as a Ptr, and that Ptr is what with_mut is called on.

When the pointee is itself a boxed value (Value<T>, i.e. Rc<RefCell<T>>), the closure takes &mut Value<T> and an extra borrow is inserted. This is the case for nested containers: the refcount model translates std::vector<std::vector<int>> as Vec<Value<Vec<i32>>> so that each element has interior mutability of its own, and a Ptr to an inner vector therefore points at a Value<Vec<i32>>, not a Vec<i32>:

#![allow(unused)]
fn main() {
ptr.with_mut(|__v: &mut Value<Vec<i32>>| (*__v.borrow_mut()).push(20))
}

The closure type is built from the C++ argument’s type, not the rule’s declared parameter type.

The read-access counterpart

For read access the converter does not emit a closure. A pointer receiver whose rule parameter is a value or & type is simply dereferenced (p.read() or (*p.upgrade().deref())); conversely, if the rule declares a Ptr parameter but the argument is not a pointer, the converter inserts an as_pointer() cast or materializes a temporary.

Preprocessor-side rewrites

Two rewrites in rule-preprocessor exist to make the with_mut rewrite possible. Both apply only to &mut parameters:

  • A * deref in front of the parameter is dropped from the body, since the substituted argument is already an lvalue or pointer expression.
  • std::mem::take(&mut aN) collapses to a bare placeholder, so the converter can re-express the move against the actual argument (for a pointer that becomes std::mem::take(&mut <lvalue>) on the borrowed pointee). The spelling must be exactly this fully qualified form: mem::take or an imported take is not rewritten. The collapsed placeholder’s access is left unknown in phase 1; phase 2 resolves the std::mem::take call to a move.

Overview

Proving ownership in the presence of C++’s unrestricted aliasing is undecidable in general, so cpp2rust does not try to satisfy Rust’s borrow checker statically. Its default output, the refcount model, moves Rust’s ownership and mutability checks to run time: reference counting replaces static ownership, and dynamic borrow checks replace static mutability checks. This trades some speed for safety, and it lets every program be translated.

libcc2rs is the runtime library where those checks live: a small crate of auxiliary types and functions, such as Value<T>, Ptr<T>, and AnyPtr, that translated programs link against. Keeping this machinery in one library keeps the generated refcount code free of unsafe.

Every Rust file cpp2rust emits imports the whole crate:

#![allow(unused)]
fn main() {
extern crate libcc2rs;
use libcc2rs::*;
}

Module map

The modules fall into three groups.

The refcounted pointer model, the core of the refcount output:

  • rc: Value<T> and Ptr<T>, the refcounted stand-ins for C values and pointers.
  • cstr: string literals and the string.h memory functions over Ptr<u8> byte strings.
  • void: AnyPtr, the type-erased pointer for void *.
  • ptr_dyn: PtrDyn<dyn T>, pointers to virtual classes.
  • reinterpret: the ByteRepr trait and allocation views that let a refcounted allocation be reinterpreted at the byte level, as C pointer casts do.
  • alloc: malloc, free, realloc, and calloc over refcounted byte arrays.

Language-feature emulation, used by both models:

  • inc and dec: traits implementing the four ++/-- operator forms.
  • iterators: iteration for C++ containers that need stable iterators, with an implementation for both refcount and unsafe, and over C strings up to the null terminator.
  • fn_ptr: FnPtr, function pointers with C-style address identity.
  • va_args: VaArg and VaList, the representation of variadic calls.
  • The goto, goto_block, and switch proc macros, re-exported from libcc2rs-macros, which rewrite unstructured control flow into state machines.

The OS and libc surface:

  • io: CFile streams, the standard streams, and read/write helpers.
  • format: printf-style format string evaluation.
  • fd: a registry tying integer file descriptors to their owning objects.
  • libc shims: safe wrappers over libc APIs, one shim.rs per rule directory (files, directories, sockets, name resolution, polling, terminal control, time, and so on), compiled into the crate at build time.
  • compat: platform-specific definitions, such as the location of errno and malloc_usable_size.

Dependencies

The crate has five dependencies:

  • libcc2rs-macros provides the control-flow and derive(ByteRepr) proc macros.
  • libc and nix provide the raw and safe OS interfaces the shims wrap.
  • jiff backs the time shims.
  • sprintf backs printf-style formatting.

Reference Counting

The refcount model produces safe Rust, and rc.rs is where that safety comes from. It defines the two types every translated program is built on: Value<T>, the translation of a C++ variable, and Ptr<T>, the translation of a C++ pointer.

Values and pointers

Rust requires every value to have a single owner, known at compile time, and references to follow the borrow rules. C++ promises neither: a variable can be aliased by any number of pointers, and any of them may write. Proving ownership in the presence of such unrestricted aliasing is undecidable in general, so the refcount model does not try. Instead it moves Rust’s ownership and mutability checks from compile time to run time, trading some speed for safety: Rc counts references and checks lifetimes dynamically, and RefCell checks at each access that readers and writers do not overlap.

A C++ variable is therefore translated as a Value<T>, an alias for Rc<RefCell<T>>. Taking the address of a variable becomes a call to as_pointer, which produces a Ptr<T>:

int b = 2;
int *b_ptr = &b;
*b_ptr = 3;
#![allow(unused)]
fn main() {
let b: Value<i32> = Rc::new(RefCell::new(2));
let b_ptr: Value<Ptr<i32>> = Rc::new(RefCell::new(b.as_pointer()));
(*b_ptr.borrow()).write(3);
}

Weak references

A C++ pointer does not own what it points to, and Ptr<T> keeps that property: it holds a Weak reference to the allocation plus an element offset. Ownership stays with the variable binding for stack values and with the allocation itself for the heap. When the owner goes away, every pointer into it dangles, and the next access panics instead of reading freed memory.

The choice of weak over strong references is about destructors. C++ RAII code relies on destructors running at precise points, such as a mutex being released at the end of a scope; a strong reference held by a stray pointer could keep the object alive past that point and run its destructor late. With weak references, objects die exactly where C++ says they do, and a pointer that outlives its object dangles.

This is the central property of the model: memory bugs of the original program, such as use after free, double free, and null or out-of-bounds dereference, become panics in the translated one. Their messages carry the ub: prefix.

Pointer kinds

A Ptr<T> knows what it points into:

  • Null: the null pointer, and the default value.
  • StackSingle and HeapSingle: a single value.
  • StackArray and HeapArray: a fixed-size array.
  • StackVec and HeapVec: a growable buffer; std::vector contents and string literals live in one.
  • Field: a field of a struct (see Pointers to fields).
  • Reinterpreted: a byte-level view produced by a cast (see Type Reinterpretation).

An array carries one reference counter for the whole allocation, not one per element: the pointer pairs a weak reference to the whole array with the offset of the element it points to, which keeps the memory and performance overhead of arrays low.

Two pointers compare equal when they point into the same allocation at the same byte offset, and ordering compares allocation addresses, as C++ pointer comparison does.

Pointers to fields

Fields are stored inline in their struct, so a struct and all of its fields share one allocation and one RefCell. A pointer to a field is to a struct what a pointer to an element is to an array: it records a weak reference to the allocation that holds the struct, its root, together with the byte offset of the field in it. field_ptr! creates one:

struct point { int x; int y; };
struct point p;
int *y = &p.y;
#![allow(unused)]
fn main() {
let p: Value<point> = Rc::new(RefCell::new(<point>::default()));
let y: Value<Ptr<i32>> = Rc::new(RefCell::new(field_ptr!(p, y)));
}

field_ptr!(p, y) works on a Value or a Ptr to a struct. The offsets are those of the C layout of the struct, which the code generator gets from Clang and writes on the fields as #[offset(N)] attributes (any constant expression works, e.g., offset_of! in the libc shims); #[derive(Record)] turns them into the implementation of the Record trait:

#![allow(unused)]
fn main() {
#[derive(Clone, Record, Default)]
pub struct point {
    #[offset(0)]
    pub x: i32,
    #[offset(4)]
    pub y: i32,
}
}

Creating a field pointer allocates nothing. The byte offset of the field is kept in the pointer’s offset, which for a field pointer counts bytes, as for a reinterpreted pointer. Pointers to fields of fields, and to fields of array elements, add up the offsets, and keep the allocation of the outermost struct or array as their root.

An access through a field pointer borrows the root, and finds the field from its offset and its type with Record::locate, which the derive generates as a comparison of the offset against those of the fields, recursing into nested structs. A field is identified by its type too, as a struct and its first field share an offset. An offset that doesn’t find a field of the pointer’s type panics with ub: invalid field pointer.

Arrays and vectors are not stored inline: an array field is a Value<Box<[T]>> of its own, and a std::vector field a Value<Vec<T>>. Pointers to their elements are ordinary array pointers, and pointers to the fields of those have the array or vector as their root. A field pointer hence always points to a single object, which is why it needs no element index.

array_field_ptr!(p, arr) is a pointer to element 0 of the array field arr, and is how the elements of an array field are accessed. It is an ordinary array pointer into the array’s Value, except when p is a reinterpreted pointer: the array then lives in the bytes of the original allocation, so the result is a reinterpreted pointer to those bytes, at the offset of the field. Writes to the elements hence reach the original allocation, and the pointer can go past the end of the struct, as C code does with a trailing char name[1] in an over-allocated struct.

Reading a field through a pointer borrows the struct only for the duration of a closure, e.g., p.with(|s| s.y), so that the borrow ends before the rest of the statement runs. field!(p, y) is the place of the field, which is read and written through p, projected to the field by a closure, without looking the field up: field!(p, y).write(3). It holds a reference to p, and takes no more room than a pointer. field_ptr! is only used where an actual pointer to the field is needed.

A field of a reinterpreted struct is itself a reinterpreted pointer, to the bytes of the field. Conversely, reinterpreting a field pointer views the bytes of the field alone, as if it were its own allocation.

Since all fields share the borrow of their struct, an expression must not write to a field while another field of the same struct is borrowed. The code generator scopes the borrows of field reads so that they end before any write in the same statement.

The heap

new and new[] are translated as Ptr::alloc and Ptr::alloc_array, and malloc, calloc, realloc, and strdup allocate through Ptr::alloc_array as well; the alloc module defines them as named functions (malloc_refcount, free_refcount, realloc_refcount, calloc_refcount, strdup_refcount, and their _unsafe twins for the unsafe model). The allocation’s Rc is leaked with Rc::into_raw so the object outlives the statement that created it, and delete and delete_array recover the leaked reference and drop it:

int *d = new int(0);
*d = 5;
delete d;
#![allow(unused)]
fn main() {
let d: Value<Ptr<i32>> = Rc::new(RefCell::new(Ptr::alloc(0)));
(*d.borrow()).write(5);
(*d.borrow()).delete();
}

delete checks that the pointer still points at the start of a live heap allocation: freeing twice, freeing through an offset pointer, or freeing a stack or Vec pointer panics with ub:.

A heap allocation can also change hands instead of being freed. to_owned_opt recovers the leaked reference the same way delete does, but returns it to the caller as an owning Option<Value<T>> (or Option<Value<Box<[T]>>> for an array), with None for the null pointer; from then on the allocation lives exactly as long as that binding. It panics for stack, Vec, and reinterpreted pointers. This is how std::unique_ptr<T> is translated: the smart pointer is an Option<Value<T>>, its constructor and reset adopt a raw pointer with to_owned_opt, and as_pointer, which is also implemented for Option<Value<T>> and yields null for None, stands in for get():

std::unique_ptr<int> u(new int(1));
int *raw = u.get();
#![allow(unused)]
fn main() {
let u: Value<Option<Value<i32>>> =
    Rc::new(RefCell::new(Ptr::alloc(1).to_owned_opt()));
let raw: Value<Ptr<i32>> = Rc::new(RefCell::new((*u.borrow()).as_pointer()));
}

Dereferences

A dereference becomes a short-lived borrow. read and write copy a value out of or into the allocation:

*d = 5;
int v = *d;
#![allow(unused)]
fn main() {
(*d.borrow()).write(5);
let v: Value<i32> = Rc::new(RefCell::new((*d.borrow()).read()));
}

A Ptr cannot simply return a &T or &mut T to its pointee: the reference would keep the RefCell borrowed with nothing to bound its lifetime. with and with_mut invert the control instead: the expression that needs the reference moves into a closure, and the borrow lasts exactly as long as the closure runs. They carry the operations that need a reference to the existing value, such as a push_back on a vector reached through a pointer (write could only replace the vector wholesale):

#![allow(unused)]
fn main() {
v.with_mut(|v| v.push(20));
}

Applied rule bodies are the main producer of these calls (see Rule Rewriting). write itself is a thin wrapper: it is defined as with_mut(|v| *v = value).

with_slice and with_slice_mut are the same idea over a range of elements: they expose len bytes starting at the pointer as a Rust slice for the duration of a closure, which is how a C buffer is passed to Rust and nix functions such as read and write.

In every case the RefCell is borrowed only for the duration of the access, which is what lets freely aliasing C++ pointers coexist with the borrow checker: no borrow outlives the expression that created it. When an expression needs an actual Rust reference, the pointer is upgraded to a StrongPtr, which holds the allocation alive and hands out a Ref.

These borrows are the model’s mutability checks, moved from compile time to run time. Rust’s rule still holds, any number of readers or one writer, but it is enforced when the access happens: an expression that writes a variable while also reading it through an alias, such as *x.borrow_mut() = *x.borrow() + 1, traps. The code generator is responsible for not emitting such expressions: it stores intermediate results in temporaries, so the reading borrow ends before the writing borrow starts.

Strong pointers

upgrade turns a Ptr<T> into a StrongPtr<T>, the same pointer holding a strong Rc to its allocation instead of a weak one:

#![allow(unused)]
fn main() {
pub enum StrongPtr<T> {
    StackSingle(Rc<RefCell<T>>),
    Vec { rc: Rc<RefCell<Vec<T>>>, offset: usize },
    StackArray { rc: Rc<RefCell<Box<[T]>>>, offset: usize },
    Field { root: Rc<dyn Root>, offset: usize },
    Reinterpreted { alloc: OriginalAlloc, byte_offset: usize, cell: RefCell<Option<T>> },
}
}

deref returns a Ref<'_, T> to the pointee. The Ref borrows the StrongPtr, so the borrow of the RefCell lasts as long as the strong pointer does: in p.upgrade().deref().field, the temporary StrongPtr lives until the end of the enclosing statement, and so does the borrow. The code generator prefers the with and with_mut closures, and upgrades only where the result of an access must borrow the pointee beyond a closure. There is no deref_mut; writes go through write and with_mut, including those to a field of a struct reached through a pointer: field!(p, x).write(v).

For the Reinterpreted variant there is no value to reference, only bytes in another allocation. deref reads those bytes into a local cell and hands out a Ref to that copy, refreshing it on every call. with_mut writes the bytes through to the original allocation before it returns, so that a write is visible through every other pointer at once.

Warning

StrongPtr is set to be removed. Holding a strong reference, even briefly, undermines the model in two ways:

  1. Nothing prevents a StrongPtr from outliving its statement. One that is stored, returned, or bound to a local keeps the object alive past the point where C++ destroys it, so its destructor runs late and dangling accesses go unnoticed; a heap object held this way makes the later delete panic. The generator only ever emits it as a temporary, but hand-edited code or a rule can break that.
  2. Even as a temporary, it lives for the whole statement. In (*p.upgrade().deref()).method() the strong reference is alive during the call, so a method that runs delete this, or otherwise deletes the object it was called on, hits delete’s reference-count check and panics with a spurious ub: invalid delete.

Where the code generator needs a different view of the same allocation, it does not upgrade at all. decay turns a pointer to a whole Vec<T> or Box<[T]> into a pointer to its first element by re-tagging the existing weak reference, and Ptr::to_dyn (see Virtual Classes) does the same for the upcast to a trait object.

Arithmetic

The offset lives in the pointer, so arithmetic never touches the allocation. p + n, p - n, and the ++/-- forms move the offset, including past the end of the allocation, exactly as C++ allows; bounds are checked only when the pointer is dereferenced. Subtracting two pointers yields their element distance and requires both to point into the same allocation.

The pointer also knows the extent of its allocation: len is the number of elements in it, whatever the pointer’s offset, and get_offset is the pointer’s element index within it. The container rules build end() and back() pointers from these (to_end, to_last) and turn a [first, last) range into a count or an absolute index the same way.

Integer casts

Casts between pointers and integers are translated as to_int and from_int:

uintptr_t n = (uintptr_t)p;
int *q = (int *)n;
#![allow(unused)]
fn main() {
let n: Value<usize> = Rc::new(RefCell::new((*p.borrow()).to_int()));
let q: Value<Ptr<i32>> =
    Rc::new(RefCell::new(<Ptr<i32>>::from_int(*n.borrow())));
}

Both currently panic when executed. Giving them well-defined semantics is work in progress (#225).

C Strings

C and C++ strings are byte strings: programs manipulate individual bytes and the contents need not be valid UTF-8, so strings are translated as u8 buffers rather than Rust String values. A string literal becomes a per-thread interned buffer with a trailing zero byte, handed out as a Ptr<u8> by Ptr::from_string_literal. Ptr<u8> also carries the memory functions C strings rely on: memcpy (with memmove semantics for overlapping buffers instead of undefined behavior), memset, memcmp, and to_rust_string for crossing into Rust APIs.

CStringIterator, returned by to_c_string_iterator, walks the bytes of a Ptr<u8> up to the null terminator; the string.h rules are built on it, and Display for Ptr<u8> prints it, so a C string can be formatted directly.

Copying a string does not go through the iterator, which would grow its destination as it goes. with_c_str scans for the null terminator in the backing storage and lends the bytes to a closure, so measuring or borrowing a string allocates nothing (c_str_len, and count on a CStringIterator, are built on it). to_c_bytes and to_rust_string, and through them strdup and the std::string constructors, copy the string with a single allocation. to_c_bytes leaves room for one more byte, which callers usually spend on the terminator or a newline.

Iterating a Ptr<T> reports how many elements are left, so collecting a Ptr into a Vec also allocates once.

void Pointers

void * is translated as AnyPtr, a type-erased Ptr. to_any erases the element type and reinterpret_cast recovers it:

char data[] = "hi";
void *vp = data;
char *cp = vp;
#![allow(unused)]
fn main() {
let data: Value<Box<[u8]>> = Rc::new(RefCell::new(Box::from(*b"hi\0")));
let vp: Value<AnyPtr> =
    Rc::new(RefCell::new((data.as_pointer() as Ptr<u8>).to_any()));
let cp: Value<Ptr<u8>> =
    Rc::new(RefCell::new((*vp.borrow()).reinterpret_cast::<u8>()));
}

reinterpret_cast returns the original pointer when the requested type matches the erased one, and a byte-level view otherwise, because C code commonly casts A * to void * and reads it back as B *.

The malloc family allocates and frees through AnyPtr, so the returned pointer is cast to the requested type and cast back to free it:

int *p = malloc(sizeof(int));
*p = 42;
free(p);
#![allow(unused)]
fn main() {
// malloc_refcount(n) is
// Ptr::alloc_array(vec![0u8; n].into_boxed_slice()).to_any()
let p: Value<Ptr<i32>> = Rc::new(RefCell::new(
    malloc_refcount(::std::mem::size_of::<i32>()).reinterpret_cast::<i32>(),
));
(*p.borrow()).write(42);
free_refcount((*p.borrow()).to_any());
}

AnyPtr also carries memcpy, memset, and memcmp, forwarding to the Ptr<u8> versions from C Strings over the byte view of its pointee.

Warning

Two AnyPtr values are equal only when they were erased from the same pointer type and compare equal as that type. A void * obtained from a Ptr<i32> and one obtained from a Ptr<u8> into the same allocation compare unequal, where C would consider them the same address. This is set to be fixed by comparing through the byte view instead.

Casts between AnyPtr and integers use the same to_int and from_int as Ptr<T>.

Virtual Classes

A pointer to a virtual class cannot be a Ptr<T>. Ptr<T> requires T: Sized, the implicit bound on every generic parameter, and a virtual class is translated as a Rust trait, whose trait object dyn T is unsized and cannot satisfy that bound. The runtime provides a dedicated PtrDyn<dyn T> type, declared with T: ?Sized, for these pointers, kept separate so the generic Ptr pays no cost for dynamic dispatch.

A PtrDyn is created at the point where C++ converts a derived pointer to a base pointer. Ptr::to_dyn takes the weak reference out of the Ptr<Derived> and applies Rust’s unsized coercion to it, turning a Weak<RefCell<Derived>> into a Weak<RefCell<dyn Base>>, without ever upgrading it. The coercion itself is written by the code generator as the closure |w| w, whose return type selects the target trait object.

Note

This coercion is why the pointee cell is a plain RefCell behind Weak rather than a struct of the runtime’s own: Weak already implements it, and a new type could only opt in through the nightly-only CoerceUnsized trait.

The conversion looks like this:

struct Base { virtual int f() const = 0; };
struct Derived : Base { int f() const override { return 1; } };

Derived d;
Base *b = &d;
int r = b->f();
#![allow(unused)]
fn main() {
let d: Value<Derived> = Rc::new(RefCell::new(<Derived>::default()));
let b: Value<PtrDyn<dyn Base>> = Rc::new(RefCell::new(
    (d.as_pointer()).to_dyn::<dyn Base>(|w| w),
));
let r: Value<i32> = Rc::new(RefCell::new(
    ({ (*(*b.borrow()).upgrade().deref()).f() }),
));
}

A virtual call goes through upgrade, which returns a StrongPtrDyn<dyn T> holding the strong reference; its deref and deref_mut borrow the object and the call dispatches through the trait’s vtable.

Warning

StrongPtrDyn is set to be removed for the same reasons as StrongPtr: it holds a strong reference that can outlive the object’s C++ lifetime, and even as a temporary it spans the whole virtual call, so a method that deletes its own object panics on delete.

PtrDyn is far smaller than Ptr: it is either null or a weak reference to a single object, on the stack or on the heap. Ptr::to_dyn keeps that distinction, so a base pointer made from a newed object can be deleted and one made from a local cannot. It has no arithmetic, no comparison, no array kinds, and no byte view. Because to_dyn is only defined for single-value pointers, a base pointer into an array of polymorphic objects (a Derived arr[N] walked through a Base *) cannot be formed.

Type Reinterpretation

C code reads the same memory at different types: a long is inspected byte by byte through a char *, a byte buffer from malloc is used as an array of structs, a struct sockaddr_in is passed where a struct sockaddr is expected. In the refcount model there are no raw bytes to point at: values are typed Rust data behind refcounted cells. The reinterpret module supplies the byte view these programs expect.

ByteRepr

ByteRepr gives a type its C byte representation:

#![allow(unused)]
fn main() {
pub trait ByteRepr: 'static {
    fn byte_size() -> usize;
    fn to_bytes(&self, buf: &mut [u8]);
    fn from_bytes(buf: &[u8]) -> Self;
}
}

The code generator emits the ByteRepr implementation of a C struct next to it:

struct header {
  int tag;
  int size;
};
#![allow(unused)]
fn main() {
pub struct header {
    pub tag: i32,
    pub size: i32,
}

impl ByteRepr for header {
    fn byte_size() -> usize {
        8
    }
    fn to_bytes(&self, buf: &mut [u8]) {
        self.tag.to_bytes(&mut buf[0..4]);
        self.size.to_bytes(&mut buf[4..8]);
    }
    fn from_bytes(buf: &[u8]) -> Self {
        Self {
            tag: <i32>::from_bytes(&buf[0..4]),
            size: <i32>::from_bytes(&buf[4..8]),
        }
    }
}
}

byte_size is sizeof(struct header). to_bytes writes each field at its C offset into an 8-byte buffer, so the buffer holds the struct exactly as it would sit in C memory. from_bytes reads such a buffer back into a fresh struct.

The primitive types serialize to their native-endian bytes, matching what C sees on the host. The libc shims implement the trait by hand with the byte layout of their C structs. Types with no meaningful C layout, such as std::fs::File or Vec<T>, implement the trait with defaults that panic, so reinterpreting one is caught at run time.

derive(ByteRepr)

libcc2rs-macros provides a #[derive(ByteRepr)] proc macro. It is implemented only for unit structs, and expanding it on a struct with fields, an enum, or a union is a compile-time error. The expansion sets byte_size to 1 and leaves to_bytes and from_bytes on the trait’s panicking defaults:

#![allow(unused)]
fn main() {
#[derive(Default, Clone, Copy, ByteRepr)]
pub struct UnitStruct;
}

Views over the original allocation

reinterpret_cast copies nothing. It produces a Ptr in the Reinterpreted kind: a handle to the original allocation plus a byte offset, stepping by the target type’s size. A read serializes the overlapping elements of the original into bytes and parses the target value out of them; a write is a read-modify-write back into the original. The data always lives in the original allocation, so writes through the original are visible through every view and writes through a view are visible everywhere else:

#![allow(unused)]
fn main() {
let p: Ptr<u64> = Ptr::alloc(0x0807060504030201);
// A view over p's allocation at byte offset 0, stepping by 1 byte.
let bytes: Ptr<u8> = p.reinterpret_cast::<u8>();

// Read: p.to_bytes() gives the 8 bytes, u8::from_bytes parses byte 0.
assert_eq!(bytes.read(), 0x01);
// Write: p.to_bytes(), replace byte 7 with 0xAA, u64::from_bytes back into p.
bytes.offset(7).write(0xAA);
// The write went into the original allocation.
assert_eq!(p.read(), 0xAA07060504030201);
}

A reinterpreted pointer counts its offset in bytes, so its arithmetic matches the C cast exactly. Casting a view again does not stack views: the new pointer keeps the handle to the original allocation.

Deleting through a reinterpreted pointer frees the original allocation. That is how free works on a buffer that has been cast around: the pointer is reinterpreted to bytes and the original allocation is deleted.

Cost

A cast allocates exactly one small object, the view, which holds a weak reference to the original allocation, the size of the target type, and a reference to a stateless table of the byte-level operations for the original’s storage type (a single value, a Vec, a boxed slice, or a field of a struct, whose bytes are found through the struct’s allocation as for a field pointer). Copying or offsetting the resulting Ptr only bumps a reference count, so a loop over a malloced array pays for the cast once, not per access.

Accessing memory through a view does not allocate: the bytes of the accessed element are staged in a stack buffer (a heap buffer is used only for accesses larger than 64 bytes). When the original allocation is a u8 buffer, as it is for everything that comes from malloc, reads and writes copy the bytes directly instead of serializing the elements one by one.

Known limitations

Reading a struct through a reinterpreted pointer builds a fresh struct with from_bytes, which exists only for the duration of the with closure, or as long as the StrongPtr that holds it. Writes to its fields go through with_mut, which encodes the struct back into the original allocation before it returns, so they are visible right away through any other pointer; a pointer to one of its fields is a reinterpreted pointer into the original allocation. However, union accessors return pointers to the union’s storage; on a reinterpreted union that storage is the temporary, so the returned pointer dangles. This is set to be fixed in the near future.

AnyPtr casts

AnyPtr::reinterpret_cast first tries to recover the pointer as it was erased: casting a void * back to the type it came from returns the original Ptr<T>, with no byte view involved. Only a cast to a different type goes through the byte representation:

#![allow(unused)]
fn main() {
let p: Ptr<u64> = Ptr::alloc(0x0807060504030201);
let any: AnyPtr = p.to_any();

// Same type as erased: the original Ptr<u64> comes back.
let back: Ptr<u64> = any.reinterpret_cast::<u64>();
assert!(back == p);

// Different type: a byte view over p's allocation, as with
// Ptr::reinterpret_cast.
let bytes: Ptr<u8> = any.reinterpret_cast::<u8>();
assert_eq!(bytes.read(), 0x01);
}

Increment and Decrement

C’s ++ and -- are expressions: x++ yields the old value and ++x the new one, and both can appear inside a larger expression. Rust only has the x += 1 statement, so the inc and dec modules define one trait per operator form:

#![allow(unused)]
fn main() {
pub trait PostfixInc { fn postfix_inc(&mut self) -> Self; }
pub trait PrefixInc  { fn prefix_inc(&mut self) -> Self; }
pub trait PostfixDec { fn postfix_dec(&mut self) -> Self; }
pub trait PrefixDec  { fn prefix_dec(&mut self) -> Self; }
}

Each method updates the value in place and returns what the C expression evaluates to: the postfix forms return a copy of the old value, the prefix forms the new one.

int x = 0;
while (x++ < 100 && x != 50) {
  ++x;
}
#![allow(unused)]
fn main() {
let x: Value<i32> = Rc::new(RefCell::new(0));
while (*x.borrow_mut()).postfix_inc() < 100 && *x.borrow() != 50 {
    (*x.borrow_mut()).prefix_inc();
}
}

The traits are implemented for the integer types with wrapping arithmetic, so overflow behaves as C’s unsigned wraparound and never panics, and for f32 and f64. Ptr<T> implements them by moving its offset one element, and the map iterators by stepping to the neighbouring key. For each translated enum the code generator emits impl_enum_inc_dec!, a macro exported by inc that implements the four traits by converting through i32.

The unsafe model uses the same traits for integers and floats. For raw pointers the same method names come from separate Unsafe* traits (UnsafePrefixInc and so on), whose methods are unsafe fn and step the pointer with offset(1), so the generated code reads the same in both models:

#![allow(unused)]
fn main() {
let mut q: *mut i32 = p;
q.prefix_inc();
q.postfix_dec();
}

Iterators

A C++ iterator is a pointer-like object: it is dereferenced, compared against end(), and moved with ++ and --, and it stays usable across the statements of a loop body. Rust iterators are consumed by a for loop and cannot be compared or stepped backwards, so the runtime represents C++ iterators with values of its own.

Random access iterators

For std::vector, std::string, and arrays the iterator is a Ptr<T> into the container’s buffer: begin() is as_pointer(), end() is to_end(), and comparison and arithmetic are the pointer’s own. Ptr<T> also implements Iterator, yielding a pointer to each element, so a range-based for becomes a Rust for over the pointer:

std::vector<int> v;
for (auto x : v)
  printf("%d\n", x);
#![allow(unused)]
fn main() {
let v: Value<Vec<i32>> = Rc::new(RefCell::new(Vec::new()));
for x in v.as_pointer() as Ptr<i32> {
    println!("{}", x.read());
}
}

Two variants serve special cases. StringIterator, returned by to_string_iterator, stops before the trailing zero byte, so iterating a std::string visits its characters only. PtrValueIter yields copies of the elements instead of pointers to them; rule bodies use it to feed a range of C memory to Rust iterator adaptors:

#![allow(unused)]
fn main() {
// std::accumulate(first, last, init)
let count = (last - first) as usize;
PtrValueIter::new(&first, count).fold(init, |acc, x| acc + x)
}

Stable iterators

std::map<K, V> is translated as a BTreeMap<K, Value<V>>, which has no addressable elements to point into. The runtime defines MapIter for it: a pair of a handle to the map and the current key, with None standing for end(). Because it stores a key rather than a position, it survives insertions and removals elsewhere in the map, as C++ guarantees. begin, end, and find_key construct one; inc and dec move to the neighbouring key; erase removes the current entry and returns the iterator to the next; the ++/-- traits and Iterator are implemented on top of these. Two iterators compare equal when they hold the same key, so it != m.end() compares Some(key) against None:

std::map<int, double> m;
for (const auto &i : m)
  sum += i.second;
#![allow(unused)]
fn main() {
let m: Value<BTreeMap<i32, Value<f64>>> =
    Rc::new(RefCell::new(BTreeMap::new()));
for i in RefcountMapIter::begin(m.as_pointer()) {
    (*sum.borrow_mut()) += (*i.second().borrow());
}
}

Warning

Equality only looks at the key: iterators into two different maps compare equal when they hold the same key. In C++ comparing them is undefined behaviour, so the translation should panic with ub: instead; this will be fixed by comparing the map handles as well.

first() and second() come from the MapIterator trait and take the place of it->first and it->second. MapIter is generic over how the map is reached, which is what gives it an implementation for both models: RefcountMapIter<K, V> holds a Ptr<BTreeMap<K, Value<V>>> and returns Value<K> and Value<V>; UnsafeMapIterator<K, V> holds a *const BTreeMap<K, Box<V>> and returns *const K and *mut V.

Function Pointers

A C function pointer can be null, compared for equality, cast to another function pointer type and back, and stored in a void *. A Rust fn value can be called and compared, but it is never null and its type is fixed, so the refcount model translates function pointers as FnPtr<T>, where T is the Rust fn type of the target:

#![allow(unused)]
fn main() {
pub struct FnPtr<T> { /* the function as first stored, and its current cast */ }

impl<T> FnPtr<T> {
    pub fn null() -> Self;
    pub fn new(f: T) -> Self;
    pub fn is_null(&self) -> bool;
    pub fn cast<U>(&self, adapter: Option<U>) -> FnPtr<U>;
    pub fn to_any(&self) -> AnyPtr;
}
}

FnPtr dereferences to the function, so a call through it is (*fp)(args). Calling a null pointer panics with ub:.

FnPtr stores the function inline, together with its address, which is how pointers are compared; the FnAddr trait provides the address. Creating, copying, and calling a function pointer does not allocate. Rust has no way to write an impl for every fn arity at once, so FnAddr is implemented by a macro for fn types of zero to sixteen parameters. A function with more parameters cannot be wrapped in an FnPtr, and taking its address fails to compile with a missing FnAddr bound.

typedef int (*int_fn)(int);
int double_it(int x) { return x * 2; }

int_fn fn = double_it;
int r = fn(5);
#![allow(unused)]
fn main() {
let fn_: Value<FnPtr<fn(i32) -> i32>> =
    Rc::new(RefCell::new(FnPtr::<fn(i32) -> i32>::new(double_it_0)));
let r: Value<i32> = Rc::new(RefCell::new((*(*fn_.borrow()))(5)));
}

Casts

C code casts function pointers to a different type and calls through the new type. When the two types are not compatible this is undefined behavior, but the argument types involved usually have the same representation, so implementations accept the call and programs rely on it. Below, add_offset takes an int *, but is called through a pointer that takes a void *:

typedef int (*generic_int_fn)(void *, int);
int add_offset(int *base, int offset) { return *base + offset; }

generic_int_fn gfn = (generic_int_fn)add_offset;
int result = gfn(&val, 42);

In Rust fn(Ptr<i32>, i32) -> i32 and fn(AnyPtr, i32) -> i32 are unrelated types, so the code generator emits an adapter: a function of the target type that converts the arguments and calls the original. cast stores it, and calls through the cast pointer go through the adapter:

#![allow(unused)]
fn main() {
let gfn: Value<FnPtr<fn(AnyPtr, i32) -> i32>> = Rc::new(RefCell::new(
    FnPtr::<fn(Ptr<i32>, i32) -> i32>::new(add_offset_4)
        .cast::<fn(AnyPtr, i32) -> i32>(Some(
            (|a0: AnyPtr, a1: i32| -> i32 {
                add_offset_4(a0.reinterpret_cast::<i32>(), a1)
            }) as fn(AnyPtr, i32) -> i32,
        )),
));
let result: Value<i32> = Rc::new(RefCell::new(
    (*(*gfn.borrow()))(val.as_pointer().to_any(), 42),
));
}

The code generator can build an adapter when the arguments and return type of the two function types have the same representation. Otherwise it passes None, and calling through the cast pointer panics with ub:.

A cast to a different type is the only operation that allocates: the pointer then also keeps the function it was created with, type-erased, so that casting back to that type can restore it. Equality compares the address of the function the pointer was created with.

Casting a function pointer to void * is to_any, and AnyPtr::cast_fn::<T> recovers it. reinterpret_cast on an AnyPtr holding a function currently panics, as do integer casts on a Ptr; both are set to be fixed in the near future.

Lambdas

A lambda is an FnPtr too, in both models (see Lambdas). One without captures is built with new from a closure, like a function. One with captures is built by the lambda! and lambda_unsafe! macros, which declare a struct holding the captures, with the body as its method, and pass both to from_lambda or from_lambda_unsafe:

#![allow(unused)]
fn main() {
impl<A, R> FnPtr<fn(A) -> R> {
    pub fn from_lambda<L>(lambda: L, call: fn(&L, A) -> R) -> Self;
    pub fn from_lambda_unsafe<L>(lambda: L, call: fn(&mut L, A) -> R) -> Self;
}
}

The unsafe model translates function pointers as Option<unsafe fn> and uses FnPtr only for lambdas. For this, FnPtrArg is also implemented for raw pointers and for Option<unsafe fn>, and derived by the structs and unions of the unsafe model.

Variadic Functions

Rust has no ... parameters and no va_list. A variadic C function is translated as a function whose last parameter is a slice of VaArg, an enum with one variant per kind of value C’s default argument promotions can produce:

#![allow(unused)]
fn main() {
pub enum VaArg {
    Int(i32),
    UInt(u32),
    Long(i64),
    ULong(u64),
    Double(f64),
    RawPtr(*mut c_void),
    Ptr(AnyPtr),
}
}

At a call site every extra argument is converted with .into(), which performs the promotions (char and short to int, float to double) and erases pointers to AnyPtr in the refcount model or *mut c_void in the unsafe model. Inside the function, va_list is a VaList, a cursor over the slice: va_start becomes VaList::new(__args), va_arg(ap, T) becomes ap.arg::<T>(), va_copy is a plain copy of the cursor, and va_end is a no-op:

int sum(int count, ...) {
  va_list ap;
  va_start(ap, count);
  int total = 0;
  for (int i = 0; i < count; i++)
    total += va_arg(ap, int);
  va_end(ap);
  return total;
}

sum(3, 10, 20, 30);
#![allow(unused)]
fn main() {
pub fn sum_0(count: i32, __args: &[VaArg]) -> i32 {
    let ap: Value<VaList> = Rc::new(RefCell::new(VaList::default()));
    (*ap.borrow_mut()) = VaList::new(__args);
    let total: Value<i32> = Rc::new(RefCell::new(0));
    // ...
    (*total.borrow_mut()) += (*ap.borrow_mut()).arg::<i32>();
    // ...
}

sum_0(3, &[10.into(), 20.into(), 30.into()]);
}

arg::<T>() goes through the VaArgGet trait, implemented for the integer and floating types, raw pointers, Ptr<T>, AnyPtr, and FnPtr<T>. Integer variants convert freely among the integer types, as va_arg does with types of the same rank; asking for a pointer where an integer was passed, or the reverse, panics, as does reading past the last argument.

Variadic libc functions such as printf and fcntl are handled by variadic rules, whose bodies receive the same &[VaArg] slice; format_c in the format module consumes one to evaluate a format string.

Control Flow Macros

Rust has no goto, and a match arm never falls into the next one. The libcc2rs-macros crate provides two procedural macros, re-exported by libcc2rs, that express these C constructs as a state machine. Both models use them.

goto_block

goto_block! takes a sequence of labeled blocks. Execution starts in the first block and falls through from each block into the next; goto!('label) jumps to the block with that label, forwards or backwards:

int retry(int n) {
  int count = 0;
  int acc = 0;
again:
  count += 1;
  acc += n;
  if (count < 3)
    goto again;
  return acc;
}
#![allow(unused)]
fn main() {
pub fn retry_0(n: i32) -> i32 {
    let n: Value<i32> = Rc::new(RefCell::new(n));
    let count: Value<i32> = <Value<i32>>::default();
    let acc: Value<i32> = <Value<i32>>::default();
    goto_block!({
        'entry: {
            *count.borrow_mut() = 0;
            *acc.borrow_mut() = 0;
        }
        'again: {
            (*count.borrow_mut()) += 1;
            (*acc.borrow_mut()) += (*n.borrow());
            if *count.borrow() < 3 {
                goto!('again);
            }
            return (*acc.borrow());
        }
    });
    panic!("ub: non-void function does not return a value")
}
}

The code generator puts the statements that precede the first C label in an 'entry block. The panic! after the block is there for the Rust compiler: the function returns from inside the state machine, but the compiler cannot see that every path does, so without a final diverging statement it rejects the function for not returning a value.

The macro expands to a loop over a match on a state variable, one arm per block. Each arm ends by setting the next state and continuing the loop, and goto!('label) sets the target state instead. In outline, the block above becomes:

#![allow(unused)]
fn main() {
let mut state: u32 = 0;
'sm: loop {
    match state {
        0 => {
            /* entry body */
            state = 1;
            continue 'sm;
        }
        1 => {
            /* again body, with goto!('again) as */
            {
                state = 1;
                continue 'sm;
            }
            break 'sm;
        }
        _ => break 'sm,
    }
}
}

break and continue written inside a block (outside any loop nested in it) still refer to the loop enclosing the goto_block!: the macro records them in a flag, leaves the state machine loop, and re-issues them after it. goto! outside a goto_block! is a compile error.

Supported goto patterns:

  • labels at the top level of a block: a function body, a loop body, or a compound statement;
  • a goto anywhere inside that block, including in nested ifs, loops, and switch cases;
  • forward and backward jumps.

Not supported yet:

  • a jump to a label that is not at the top level of a block enclosing the goto, such as from outside a loop to a label in its body;
  • a jump to a label inside an if branch.

switch

A switch without fallthrough is translated as a plain match inside a labeled block, where break becomes a break out of that block. When some case falls into the next, the code generator uses switch! instead. It is written like a match, but an arm whose body does not end in break continues into the body of the following arm, as C does:

switch (x) {
case 1:
  r += 10;
case 2:
  r += 20;
  break;
default:
  r = -1;
  break;
}
#![allow(unused)]
fn main() {
switch!(match (*x.borrow()) {
    v if v == 1 => {
        (*r.borrow_mut()) += 10;
    }
    v if v == 2 => {
        (*r.borrow_mut()) += 20;
        break;
    }
    _ => {
        (*r.borrow_mut()) = -1;
        break;
    }
});
}

switch! desugars to a goto_block! whose first block dispatches on the condition to the block of the matching case; the case bodies follow as consecutive blocks, so falling off the end of one enters the next, and break leaves the whole switch!. A continue in a case is not captured by the switch!: as in C, it continues the loop enclosing the switch, and is a compile error when there is none. goto and switch mix freely: a switch! can be nested in a goto_block!, a goto! inside a case can target a label of the enclosing block, and a label attached to a case is supported. Statements between the switch and its first case are not supported yet.

Hoisted declarations

In C a variable declared in one case is visible in the cases after it, because they all belong to the same block. Each switch! arm is a separate Rust block, so the code generator hoists such declarations above the macro and leaves an assignment in the case:

switch (x) {
case 1:
  r = 1;
  int y;
  y = 10;
  r += y;
case 2:
  y = 20;
  r = y;
  break;
}
#![allow(unused)]
fn main() {
let y: Value<i32> = <Value<i32>>::default();
switch!(match (*x.borrow()) {
    v if v == 1 => {
        (*r.borrow_mut()) = 1;
        *y.borrow_mut() = 10;
        (*r.borrow_mut()) += *y.borrow();
    }
    v if v == 2 => {
        *y.borrow_mut() = 20;
        (*r.borrow_mut()) = *y.borrow();
        break;
    }
    _ => {}
});
}

The same hoisting applies to variables used across the labeled blocks of a goto_block!, as count and acc above show.

I/O and Formatting

The io, format, and fd modules support the stdio stream functions, printf-style formatting, and descriptor-based I/O.

C Streams

A C FILE is more than a file handle: it carries sticky end-of-file and error flags that feof and ferror report long after the read that set them. std::fs::File keeps no such state, so the refcount model translates FILE * as a Ptr<CFile>, a libc shim that holds the file descriptor together with these two flags. The standard streams are thread-local CFile values over descriptors 0, 1, and 2, returned by c_stdin, c_stdout, and c_stderr. CFile does no buffering at present: every read or write on it is a system call on the descriptor.

In the unsafe model streams stay raw: stdin_unsafe, stdout_unsafe, and stderr_unsafe return the process’s *mut libc::FILE handles, whose symbol names differ per platform (stdin on Linux, __stdinp on macOS).

fread and fwrite exist in both models as named functions, because translated programs take their address:

#![allow(unused)]
fn main() {
pub fn fread_refcount(
    a0: AnyPtr,
    a1: usize,
    a2: usize,
    a3: Ptr<CFile>,
) -> usize;
pub unsafe fn fread_unsafe(
    a0: *mut c_void,
    a1: usize,
    a2: usize,
    a3: *mut libc::FILE,
) -> usize;
}

The refcount variant reinterprets the destination as a byte array and reads through the CFile; the unsafe variant forwards to libc::fread.

C++ Streams

In the refcount model cin, cout, and cerr are translated as Ptr<std::fs::File> values over duplicates of the standard descriptors, returned by the cin, cout, and cerr functions. Ptr<T> implements write_fmt and write_all whenever T: Write, forwarding to the pointee through with_mut, so cout << x becomes write!(cout(), "{}", x) and a raw byte range is written with cout().write_all(..). In the unsafe model cin_unsafe, cout_unsafe, and cerr_unsafe return raw pointers to thread-local std::fs::File values. C++ streams do not map fully onto std::fs::File, so this translation may change in the future.

Formatting

The code generator first translates the printf family into the idiomatic print! and println! macros. That is not always possible: the target stream may not be known at translation time, or the format string may be a runtime value. For those cases, and for functions that format into a buffer such as snprintf, the refcount model falls back to format_c; the unsafe model calls libc directly.

format_c evaluates a C format string against a slice of variadic arguments and returns the formatted String:

#![allow(unused)]
fn main() {
pub fn format_c(fmt: &str, va: &[VaArg]) -> String;
}

Parsing and rendering come from the sprintf crate. The integer, character, string, and floating-point conversions are supported, and %s reads the argument through the refcounted pointer as a Rust string. A malformed format string or an argument of the wrong kind is a panic. Three things are not supported yet:

  1. %p renders through the pointer’s integer cast, which currently panics.
  2. %n is not handled.
  3. A * width or precision (%*d, %.*s) is parsed but its integer argument is not consumed, so the remaining arguments are misaligned.

File descriptors

Rust tracks descriptor ownership in the type system: an OwnedFd closes the descriptor when dropped, and a BorrowedFd grants temporary access to one. C has no such distinction: a descriptor is a plain int, mixed freely with integer arithmetic, so the translator cannot tell which int values are descriptors. The refcount model therefore leaves descriptors as integers in the translated program and keeps the ownership in one place, the thread-local FdRegistry, a table from each integer to the open descriptor it names.

The registry follows the descriptor’s life. When a rule opens a file, FdRegistry::register stores the resulting OwnedFd and hands the program its raw number. When a rule performs I/O on that number, FdRegistry::with_fd looks the entry up and lends it out as a BorrowedFd for the duration of the call. When the program calls close, FdRegistry::close removes the entry, which closes the descriptor. The registry starts out holding the standard descriptors 0, 1, and 2.

In the fstat rule, the descriptor argument goes through with_fd:

#![allow(unused)]
fn main() {
fn f2(a0: i32, a1: Ptr<Stat>) -> i32 {
    match FdRegistry::with_fd(a0, |fd: BorrowedFd<'_>| {
        nix::sys::stat::fstat(fd)
    }) {
        // ...
    }
}
}

with_fds borrows several descriptors at once for select-style calls. The select rule collects every descriptor set in the fd_set arguments and borrows them all for the duration of the call:

#![allow(unused)]
fn main() {
let wanted: Vec<i32> = /* the descriptors set in the fd_set arguments */;
FdRegistry::with_fds(&wanted, |borrowed: &[BorrowedFd<'_>]| {
    let mut read_set = nix::sys::select::FdSet::new();
    for fd in &borrowed[..read_count] {
        read_set.insert(*fd);
    }
    // ... build the write and except sets the same way ...
    nix::sys::select::select(nfds, &mut read_set, /* ... */)
})
}

Using a descriptor that was never opened, or using it after it was closed, is a bug in the original program. The registry turns such a use into a panic (with a message prefixed ub:) so the bug surfaces instead of going unnoticed.

libc Shims

libc structs hold raw pointers, which are incompatible with the refcounted pointers the refcount model uses, so a libc struct cannot be used directly. The shim modules therefore define Rust counterparts for the libc types translated programs use. A shim struct mirrors its C struct member by member, like a translated struct: fields are stored inline, except for arrays, which are Value<Box<[T]>>s of their own (see Boxing). A shim converts to or from the underlying libc or nix type at the call boundary.

Stat is a typical shim:

#![allow(unused)]
fn main() {
#[derive(Clone, Default, Record)]
pub struct Stat {
    #[offset(offset_of!(::libc::stat, st_dev))]
    pub st_dev: u64,
    #[offset(offset_of!(::libc::stat, st_ino))]
    pub st_ino: u64,
    // ...
    #[offset(offset_of!(::libc::stat, st_size))]
    pub st_size: i64,
}

impl Stat {
    pub fn from_libc(s: &::libc::stat) -> Self { /* ... */ }
}
}

A stat call in the source program becomes a nix::sys::stat::stat call. On success nix returns a raw libc::stat, so the result goes through Stat::from_libc before it is written into the translated struct.

Like translated structs, shims derive Record, so that pointers to their fields can be taken (see Pointers to fields). The offsets of the fields are those of the libc struct, given by offset_of!, and their ByteRepr gives the size of the libc struct, which locates the fields of the elements of an array of shims, like an array of pollfd. The sockaddr family, whose byte representation has a layout of its own, uses the offsets of that layout.

The modules

Each shim lives next to the rules that use it, as rules/<dir>/shim.rs. The libcc2rs build script finds every such file and includes it as a module of the crate, re-exported at the crate root, so a shim refers to other runtime items through crate:: and translated code reaches it as libcc2rs::Stat.

Rule dirC types
stdioFILE (CFile)
direntstruct dirent, DIR (Dirent, CDir)
selectfd_set (CFdSet)
ifaddrsstruct ifaddrs (Ifaddrs)
ipstruct in_addr, struct in6_addr (InAddr, In6Addr)
netdbstruct addrinfo (Addrinfo)
pollstruct pollfd (Pollfd)
pwdstruct passwd (Passwd)
socketthe sockaddr family (Sockaddr, SockaddrIn, SockaddrIn6, SockaddrUn, SockaddrStorage)
statstruct stat (Stat)
termiosstruct termios, struct winsize (Termios, Winsize)
timestruct tm, struct timeval, struct timespec (Tm, Timeval, Timespec)

Most shims are plain data plus conversions like Stat. CFile carries the stdio stream logic (see I/O and Formatting), and the time shims convert through the jiff crate. CFdSet and the sockaddr family need more than a field-by-field mirror and are described in their own sections below.

Each shim file also gives the raw libc struct it mirrors an empty ByteRepr impl (impl ByteRepr for ::libc::stat {}), whose methods panic. These exist so that the generated ByteRepr implementation of a translated struct with a libc struct member still compiles; reinterpreting such a struct is not supported at present.

CFdSet

nix has its own FdSet, but it is stricter than the C one: it ties the set to the lifetimes of the descriptors it holds. A C fd_set is just a set of integers that accepts anything; whether the descriptors are valid is only checked by the select call that eventually receives the set. CFdSet keeps the C behavior by storing plain integers, and the select rule builds the nix FdSet from it at call time.

The sockaddr family

C socket code reinterprets one address struct as another: the program fills in a struct sockaddr_in, passes it to bind as a struct sockaddr *, and casts back to the concrete type on the way out of accept. The address shims keep this pattern working by implementing ByteRepr with the exact byte layout of their C structs: the family in the first two bytes, the remaining members at their C offsets. A cast in the source program becomes a reinterpret_cast on the refcounted pointer, which reads the struct through that byte layout as the target type, so any member of the family can be viewed as any other, exactly as in C.

The call boundary works the same way. Sockaddr::decode reads the family from the first two bytes and reinterprets the pointer as the concrete type before handing nix a typed address:

#![allow(unused)]
fn main() {
pub fn decode(
    addr: &Ptr<Sockaddr>,
    _len: u32,
) -> Option<Box<dyn SockaddrLike>> {
    let family = addr.reinterpret_cast::<u16>().read();
    if family == libc::AF_INET as u16 {
        let m = addr.reinterpret_cast::<SockaddrIn>().read();
        Some(Box::new(nix::sys::socket::SockaddrIn::from(m.to_libc())))
    }
    // ... AF_INET6 and AF_UNIX in the same way ...
}
}

Sockaddr::encode goes the other way, writing an address returned by nix into the caller’s buffer through the concrete shim. Ifaddrs hands out its addresses as Ptr<Sockaddr> values ready to be reinterpreted.

Non-uniform fields

Some struct fields are not spelled the same on every platform. struct stat keeps the modification time in a nested struct timespec, named st_mtim on Linux and st_mtimespec on macOS, while the shim exposes a single st_mtime field. struct in6_addr hides its bytes behind the internal __in6_u union on Linux, while the shim exposes s6_addr. The shims pick one uniform field, and the code generator meets them halfway: replaceNonUniformLibcField in the converter rewrites the platform-specific member chain in the source, so st.st_mtim.tv_sec becomes st.st_mtime in the translated code.

Compat Helpers

Some C interfaces are macros or platform-specific symbols rather than plain functions. On the source side, cpp2rust rewrites them into ordinary calls (see Compat Shims); the compat module is the runtime side of that rewrite.

errno expands to a platform-specific function call (__errno_location on Linux, __error on macOS).

In the unsafe model, cpp2rust_errno_unsafe binds both platform symbols under one name and returns the real libc errno location:

#![allow(unused)]
fn main() {
pub unsafe fn cpp2rust_errno_unsafe() -> *mut i32;
}

In the refcount model, errno is a thread-local refcounted i32 that the runtime maintains itself:

#![allow(unused)]
fn main() {
pub fn cpp2rust_errno() -> Ptr<i32>;
}

Refcount code reaches the operating system through the libc shims and nix, so libc’s errno is never read by this model. Keeping the cell current is a discipline of the rules: every rule that translates a call that can fail must write the error code into cpp2rust_errno() on the failure path (see Compat Shims); nothing enforces this, and a rule that skips the write breaks programs that check errno.

malloc_usable_size is bound under one name for both platforms (the symbol is malloc_size on macOS).

Overview

This part of the book documents the internals of the code generator: how the clang AST is traversed and how Rust code is emitted.

The Translation Pipeline

flowchart TD
    driver["<b>cpp2rust</b><br/><code>cpp2rust/cpp2rust.cpp</code>"]
    lib["<b>TranspileSrc / TranspileDir</b><br/><code>cpp2rust/cpp2rust_lib.cpp</code>"]
    action["<b>FrontendAction, ASTConsumer</b><br/><code>cpp2rust/ast_consumer.cpp</code>"]
    factory["<b>CreateConverter</b><br/><code>cpp2rust/converter/factory.cpp</code>"]
    mapper["<b>Mapper::LoadTranslationRules</b><br/><code>cpp2rust/converter/mapper.cpp</code>"]
    conv["<b>Converter / ConverterRefCount</b><br/><code>cpp2rust/converter/</code>"]
    out["<b>output file</b><br/>rustfmt"]
    driver -->|"source or compilation database"| lib
    lib -->|"one per translation unit"| action
    action --> factory
    factory -.->|"first call only"| mapper
    factory --> conv
    conv -->|"rs_code"| out

The stages

  1. cpp2rust (cpp2rust/cpp2rust.cpp) parses the flags, resolves the rules directory (see Loading and Matching), and calls TranspileSrc for --file or TranspileDir for --dir.
  2. TranspileSrc / TranspileDir (cpp2rust/cpp2rust_lib.cpp) run clang tooling over the source or over every file in compile_commands.json, with one FrontendAction per translation unit.
  3. ASTConsumer::HandleTranslationUnit calls CreateConverter (cpp2rust/converter/factory.cpp), which loads the translation rules on its first call and constructs a Converter (--model=unsafe) or a ConverterRefCount (--model=refcount).
  4. The converter emits the file preamble if this is the first unit, then traverses the unit and appends Rust text to rs_code.
  5. The driver writes rs_code to the -o path and runs rustfmt on it.

Gotchas

  • Every unit is parsed with the platform flags from cpp2rust/compat/platform_flags.h, which put the compat headers ahead of the system headers and set -D_FORTIFY_SOURCE=0, so macro-heavy libc APIs reach the converter as plain function calls.
  • In --dir mode __FILE__ is redefined to the file’s basename, so the generated code does not embed the absolute paths of the build machine.
  • Rules are loaded once per process, and the file preamble is emitted once, by the first unit; the bookkeeping that spans units (which declarations and records have already been emitted) is kept in static members of Converter.
  • After the last unit, Converter::EmitOpaqueRecords appends pub struct Name; for every record type that was referenced but never defined, so types only used behind pointers still compile.
  • A failing rustfmt is reported as an error, but the unformatted file stays on disk for inspection.

Types

Every place the converter prints a type goes through Convert(QualType). It first asks the type rules for a mapping, so library types and typedef names such as size_t are resolved by rules, and only falls back to the Visit*Type methods for the built-in and user-defined types described here.

Given

struct Item {
  int id;
  char name[8];
  std::vector<int> refs;
};

int count(Item item) { return item.id; }

the unsafe model produces (attributes and trait impls omitted)

#![allow(unused)]
fn main() {
pub struct Item {
    pub id: i32,
    pub name: [libc::c_char; 8],
    pub refs: Vec<i32>,
}
pub unsafe fn count_0(mut item: Item) -> i32 {
    return item.id;
}
}

and the refcount model produces

#![allow(unused)]
fn main() {
pub struct Item {
    #[offset(0)]
    pub id: i32,
    #[offset(4)]
    pub name: Value<Box<[u8]>>,
    #[offset(16)]
    pub refs: Value<Vec<i32>>,
}
pub fn count_0(item: Item) -> i32 {
    let item: Value<Item> = Rc::new(RefCell::new(item));
    return (*item.borrow()).id;
}
}

Fields are stored inline in their struct, so a whole struct lives in a single Value, like the elements of an array; only arrays and vectors are Values of their own (see Boxing). A pointer to a field records the allocation of the struct plus the byte offset of the field in it, which the #[offset(N)] attributes give (see Pointers to fields).

Type Mappings

The table gives the spelling of each C++ type in both models, before any refcount boxing. T stands for the translated inner type.

C++Unsafe modelRefcount model
boolboolbool
int, unsigned long, …i32, u64, … (host width)same
float, doublef32, f64same
charlibc::c_charu8
size_t and other typedefsby type rule (usize), else desugaredsame
T[N][T; N]Box<[T]>
T[][T]Box<[T]>
struct S, enum ES, Esame
T *, T &*mut T, *const TPtr<T>
Abstract **mut dyn AbstractPtrDyn<dyn Abstract>
void **mut ::libc::c_voidAnyPtr
R (*)(A)Option<unsafe fn(A) -> R>FnPtr<fn(A) -> R>
va_listVaListVaList
lambda closureimpl Fn(A) -> R as a parameter, _ elsewheresame
std::unique_ptr<T>by type rule (Option<Box<T>>)by type rule (Option<Value<T>>)
std::vector<T> and other STLby type rule (Vec<T>)by type rule (Vec<T>, Vec<Value<Vec<T>>> when nested)

Other built-ins (wchar_t, long double, char16_t) are omitted. Rvalue references (T &&) have no mapping of their own; they reach the converter only through std::move and implicit move constructors, which are handled by rules and by the constructor translation.

User-defined types as rules

When a record or enum declaration is converted, Mapper::AddRuleForUserDefinedType registers it in the mapper’s type table: the C++ name maps to the Rust name, and its pointer form maps to *mut Name or Ptr<Name> (*mut dyn Name or PtrDyn<dyn Name> for abstract classes); nested records are registered too. This is what makes library types instantiated with user types translatable: Mapper::Map matches std::vector<Item> against the rule for std::vector<T1> and then has to map T1 = Item through the same table, which would fail if Item were not in it.

Scalars

char is libc::c_char in the unsafe model, whose signedness follows the platform like C’s, and u8 in the refcount model, because C strings are byte vectors there (see C Strings). Since most C implementations have signed char, the refcount model is set to switch to i8 (#246).

Arrays

In the refcount model a constant array becomes Box<[T]>, dropping the length. Ptr<T> carries only the element type, not N, so a [T; N] could not be pointed to without a Ptr per length; Box<[T]> gives arrays of every length, and heap arrays, the same shape. [T; N] survives only inside sizeof, which becomes ::std::mem::size_of::<[T; N]>().

Array parameters decay to pointers as in C.

Typedefs and qualifiers

Typedef names are looked up as type rules before being desugared, which is how size_t maps to usize instead of the underlying unsigned long.

Constness is dropped: in the unsafe model it survives only as *const on pointers and as a missing mut on bindings, and in the refcount model it has no representation.

The Pointers and References page covers how values of pointer types are read and written.

Boxing

In the refcount model a variable is boxed: its type T is wrapped in Value<T>, an alias for Rc<RefCell<T>> (see Reference Counting). Without the box, taking the address of a variable would need a Rust reference, and arbitrary C++ aliasing cannot be expressed with references.

Not every type position is boxed. ConverterRefCount keeps a stack of conversion kinds, conversion_kind_, and the construct that owns the type pushes one before printing it:

  • FullRefCount: pushed by variable declarations; Convert(QualType) wraps the result in Value<...>.
  • Pointee: pushed by field declarations; the bare type is printed. Fields that are arrays, or whose type maps to a Vec or a Box (std::vector, std::string, std::array), push FullRefCount instead (see below).
  • Unboxed: pushed by parameter lists, return types, and record names; the bare type is printed.
  • Ptr: pushed by a pointer type for its pointee; also printed bare.

The result by position:

PositionintItemint[3]
local variable, globalValue<i32>Value<Item>Value<Box<[i32]>>
function parameter, return typei32Itemdecays to Ptr<i32>
struct fieldi32ItemValue<Box<[i32]>>
pointee of Ptr<T>, element of a containeri32ItemBox<[i32]>

Parameters arrive unboxed and are re-boxed by the function preamble; return values are unboxed:

int add(int a, Item item) { return a + item.id; }
#![allow(unused)]
fn main() {
pub fn add_0(a: i32, item: Item) -> i32 {
    let a: Value<i32> = Rc::new(RefCell::new(a));
    let item: Value<Item> = Rc::new(RefCell::new(item));
    return *a.borrow() + (*item.borrow()).id;
}
}

C++ passes arguments to functions by copy, so signatures stay unboxed; boxing the copy on entry then lets the body treat parameters exactly like local variables. The preamble skips reference parameters, which are a Ptr<T> and never boxed.

Nested containers, library ones and arrays alike, box each level except the innermost, so that every inner container can be borrowed and mutated on its own, and a pointer can be taken to it. The boxing is written into the type rules themselves: std::vector<std::vector<int>> maps to Vec<Value<Vec<i32>>>, and the carray rules map int a[2][2] to Box<[Value<Box<[i32]>>]>, both before the outer Value<...> of the declaration is added.

Struct fields are stored inline, so that a whole struct is a single allocation, and a pointer to a field records the struct’s allocation and the field’s byte offset (see Pointers). Arrays and vectors are the exception: an array field is a Value<Box<[T]>> of its own, and a vector field a Value<Vec<T>>. A pointer to an element, or to a field of an element, then has the array or the vector as its allocation instead of the struct, and pointer arithmetic moves between elements as for any other array:

struct Holder { std::vector<Point> points; int n; };
h.points[0].y = 5;
#![allow(unused)]
fn main() {
pub struct Holder {
    #[offset(0)]
    pub points: Value<Vec<Point>>,
    #[offset(24)]
    pub n: i32,
}
(*(*h.borrow()).points.borrow_mut())[(0_usize) as usize].y = 5;
}

An array or vector field is accessed like a local one, through its own borrow() or borrow_mut(), and the struct is only borrowed immutably to reach it. As Value is shared on clone(), structs with such fields implement Clone by copying the arrays and vectors, instead of deriving it.

Naming

Rust has one flat namespace per module and no overloading, so C++ names are flattened and disambiguated when they are emitted.

Records and enums are named by Mapper::ToRustName from their qualified C++ spelling: ::, <, >, commas, and spaces all become _. So ns::Foo is ns_Foo, the instantiation MyContainer<int> is MyContainer_int_, and a struct Level1 nested in Level0 is Level0_Level1. The same name is used for the struct, its impl blocks, and every mention of the type.

An anonymous struct, union, or enum is named anon_N, numbered in order of first appearance. In C, an anonymous tag that is only reachable through a typedef (typedef struct { ... } Point;) is emitted as Point_struct (or Point_enum), because C keeps tags and ordinary identifiers in separate namespaces and Point may already be a variable or function.

Names that are Rust keywords get a trailing underscore: a variable type becomes type_. The same applies to a keyword followed only by underscores, so a C++ identifier that was already type_ becomes type__ and cannot collide with the renamed type.

Free functions and global variables get a numeric suffix (main_0, foo_3) from a process-wide table keyed by mangled name, which keeps overloads and same-named static functions from different files apart. Methods keep their name unless they are overloaded, in which case the parameter types are appended (method_i32, method_i32_const). operator< is emitted as lt; comparison operators additionally produce the corresponding trait impls (PartialOrd, Ord, PartialEq).

Copy and move constructors are named copy_from and move_from, and copy and move assignment operators copy_assign and move_assign, as long as the class has only one member of that kind and no method already uses the name; otherwise they get the overloaded form.

Classes and Structs

A class becomes a struct with one field per data member, an impl block holding its constructors and methods, and trait implementations after it. Given

class Counter {
  int count_;

public:
  Counter(int start) : count_(start) {}
  ~Counter() { count_ = 0; }
  int get() const { return count_; }
  void set(int v) { count_ = v; }
};

the unsafe model produces

#![allow(unused)]
fn main() {
#[repr(C)]
#[derive(Copy, Clone, Default)]
pub struct Counter {
    count_: i32,
}
impl Counter {
    pub unsafe fn Counter(mut start: i32) -> Self {
        let mut this = Self { count_: start };
        this
    }
    pub unsafe fn get(&self) -> i32 {
        return self.count_;
    }
    pub unsafe fn set(&mut self, mut v: i32) {
        self.count_ = v;
    }
}
}

and the refcount model produces

#![allow(unused)]
fn main() {
#[derive(Default)]
pub struct Counter {
    count_: Value<i32>,
}
impl Counter {
    pub fn Counter(start: i32) -> Self {
        let start: Value<i32> = Rc::new(RefCell::new(start));
        let mut this = Self {
            count_: Rc::new(RefCell::new(*start.borrow())),
        };
        this
    }
    pub fn get(&self) -> i32 {
        return *self.count_.borrow();
    }
    pub fn set(&self, v: i32) {
        let v: Value<i32> = Rc::new(RefCell::new(v));
        *self.count_.borrow_mut() = *v.borrow();
    }
}
impl Drop for Counter {
    fn drop(&mut self) {
        *self.count_.borrow_mut() = 0;
    }
}
impl Clone for Counter {
    fn clone(&self) -> Self {
        let mut this = Self {
            count_: Rc::new(RefCell::new(*self.count_.borrow())),
        };
        this
    }
}
impl ByteRepr for Counter { /* byte_size, to_bytes, from_bytes */ }
}

The unsafe model adds #[repr(C)] and derives what it can; the refcount model writes most impls by hand. Which traits are emitted, and when they are derived rather than written, is on the Traits page.

Fields keep their C++ access: pub for public members, nothing for private ones. A class nested in another class is emitted as its own top-level struct, named Outer_Inner (see Naming); Rust has no nested types, and the outer struct refers to it by that name. A record that is only forward-declared, or whose definition is never converted because it is only used behind pointers, is emitted at the end of the file as an empty pub struct Name; (see The Translation Pipeline).

A constructor becomes an associated function named after the class. It opens with let mut this = Self { ... }, one field per member initializer, then runs the C++ body and returns this. Copy and move constructors become copy_from and move_from (see Naming); the Clone impl calls copy_from when the copy constructor is user-defined.

Methods take &self when const and &mut self otherwise in the unsafe model. In the refcount model they always take &self, since mutation goes through the fields’ RefCells. Inside a method, this is self.

A destructor with a body becomes impl Drop in the refcount model.

Warning

The unsafe model does not emit destructors at all; a user-defined destructor is silently dropped (#310).

Inheritance

An abstract class becomes a trait with one method per pure virtual function, and a class deriving from it implements the trait with its overrides. Given

class Animal {
public:
  virtual bool bark() const = 0;
};

class Dog : public Animal {
  bool bark() const override { return true; }
};

the unsafe model produces (attributes omitted)

#![allow(unused)]
fn main() {
pub unsafe trait Animal {
    unsafe fn bark(&self) -> bool;
}
pub struct Dog {}
unsafe impl Animal for Dog {
    unsafe fn bark(&self) -> bool {
        return true;
    }
}
}

and the refcount model produces (attributes and the Clone and ByteRepr impls omitted)

#![allow(unused)]
fn main() {
pub trait Animal {
    fn bark(&self) -> bool;
}
pub struct Dog {}
impl Animal for Dog {
    fn bark(&self) -> bool {
        return true;
    }
}
}

Non-virtual methods of the derived class go into its own impl Dog block as usual. Because the base is a trait, pointers to it are *mut dyn Animal in the unsafe model and PtrDyn<dyn Animal> in the refcount model, and a Dog * is upcast at the call site. Only the first base class is considered, and only virtual methods go through the trait; bases with data members or non-virtual methods, and multiple inheritance, are outside the supported subset.

Templates

Class templates are translated by full instantiation: each instantiation used by the program becomes its own struct and impl block, named after the template arguments (see Naming). MyContainer<int> and MyContainer<char> become MyContainer_int_ and MyContainer_char_, each with a complete copy of the methods specialized for its element type. Nothing is shared between instantiations, and Rust generics are not used.

Flexible array members

A trailing array member of size 0, 1, or [] that C code over-indexes into memory allocated past the struct is detected with clang’s isFlexibleArrayMemberLike. In the unsafe model an access to such a member is not an array index, which Rust would bounds-check against the declared length, but pointer arithmetic from the array’s start: s.bytes[i] becomes *s.bytes.as_mut_ptr().add(i as usize), and &s.bytes[i] the same without the leading *. The refcount model has no dedicated handling; the pattern works when the array is a union member, because the union accessor returns a Ptr over the whole allocation that can be offset freely.

Traits

Every emitted struct comes with a fixed set of trait implementations. Some are derived, some are written out; which is which depends on the model. Enums derive Clone, Copy, PartialEq, Debug, Default in both models, and unions derive Copy, Clone in the unsafe model regardless of their fields; the table below is for structs.

TraitUnsafe modelRefcount model
Copyderived when every field is copyablenever
Clonederivedderived when member-wise, else hand-written
Defaultderived when possible, else hand-writtensame
Dropnot emitted (#310)hand-written from a user destructor with a body
Ord, PartialOrd, PartialEq, Eqhand-written from operator<same
ByteReprnot neededhand-written for every record and enum
Recordnot neededderived for every struct

Copy and Clone

In the unsafe model a record derives Copy unless a field is translated to a Vec, BTreeMap, Option<Box<T>> (std::unique_ptr), or a record that is not itself Copy; Clone is derived unless the C++ copy constructor is deleted. The refcount model skips Clone altogether when the copy constructor is deleted, and never derives Copy. It derives Clone for C structs and for classes with an implicit or defaulted copy constructor, which copy each field with its own clone, i.e., its C++ copy constructor. The exceptions are fields that are or nest a Value, such as std::vector<int> (Value<Vec<i32>>, see Boxing) or std::vector<std::vector<int>> (Value<Vec<Value<Vec<i32>>>>), whose derived clone would share the Values instead of copying them; the generated Clone then translates the implicit copy constructor, which copies them deeply.

A user-defined copy constructor takes a Ptr to the source, while clone only has &self. clone passes it a pointer to a shallow copy of self, built field by field without running any copy constructor, so that the source is copied exactly once.

This is what gives struct assignment and pass-by-value C++’s member-by-member copy.

Default

Default is the value of a T x; without initializer, of T x = {}, and of the elements of new T[n]. It is derived when the derived impl gives the C zero value, and hand-written otherwise: when the class has a user-defined default constructor, default() calls it; when a field is a C array, a std::array, a function pointer, or a libc record, default() builds the struct field by field, each with the same default value the converter uses for a variable of that type declared without an initializer:

#![allow(unused)]
fn main() {
impl Default for S {
    fn default() -> Self {
        S {
            head: 0_i32,
            tail: [0_i32; 3],
            buf: [0 as libc::c_char; 4],
        }
    }
}
}

Unions always get a hand-written impl that zeroes their bytes.

Drop

A user-defined destructor with a non-empty body becomes impl Drop, with the body translated as a method body. Only the refcount model emits it; the unsafe model drops destructors silently (#310).

Comparison

A class that defines operator< (as a method or an out-of-line function) gets Ord, PartialOrd, PartialEq, and Eq, all expressed through the emitted lt method: cmp calls it both ways to pick Less, Greater, or Equal, and eq is “neither is less”. Only one comparison operator per class is supported, and only operator<. The converter assumes the operator is const, which Rust requires (cmp and eq take &self) but C++ does not; a non-const operator< is still emitted as lt(&self, ...).

ByteRepr

The refcount model emits ByteRepr for every record and enum: byte_size, to_bytes, and from_bytes laid out with the C offsets of the fields (enums go through their i32 value). It is what lets a Ptr to the type be reinterpreted as bytes, and bytes be read back as the type. A record with a field that has no byte representation gets an impl with byte_size alone, which locates the elements of arrays of the record for field pointers, and reinterpreting it panics at run time.

Record

#[derive(Record)], from libcc2rs-macros, gives pointers to fields access to the fields of a struct. The converter writes the C offset of each field, from Clang’s record layout, as an #[offset(N)] attribute on it, and the derive generates the table of offsets that field_ptr! looks fields up in, and the locate functions that find a field from its offset when a field pointer is accessed.

Enums

An enum becomes a Rust enum with one variant per enumerator and explicit discriminants, followed by the impls that give it C semantics. Given

enum Color { RED, GREEN, BLUE };

both models produce

#![allow(unused)]
fn main() {
#[derive(Clone, Copy, PartialEq, Debug, Default)]
enum Color {
    #[default]
    RED = 0,
    GREEN = 1,
    BLUE = 2,
}
impl From<i32> for Color {
    fn from(n: i32) -> Color {
        match n {
            0 => Color::RED,
            1 => Color::GREEN,
            2 => Color::BLUE,
            _ => panic!("invalid Color value: {}", n),
        }
    }
}
libcc2rs::impl_enum_inc_dec!(Color);
}

and the refcount model adds an impl ByteRepr for Color that converts through the i32 value.

The first enumerator is the #[default], which is what a zero-initialized or default-constructed enum variable holds. From<i32> is the integer-to-enum cast; a value that matches no enumerator panics, where C would silently keep the integer. impl_enum_inc_dec! implements the four ++/-- forms by stepping through the enumerators. Enum-to-integer casts need no impl and become as i32.

enum class is translated the same way; enumerators are always spelled Color::RED on the Rust side, scoped or not. In C, an anonymous enum named only through a typedef (typedef enum { ... } Tag;) is emitted as Tag_enum, and an anonymous enum with no name at all as anon_N (see Naming).

Unions

Given

union Number {
  int i;
  float f;
};

int foo(void) {
  union Number u;
  u.i = 42;
  return u.i;
}

the unsafe model produces

#![allow(unused)]
fn main() {
#[repr(C)]
#[derive(Copy, Clone)]
pub union Number {
    pub i: i32,
    pub f: f32,
}
impl Default for Number {
    fn default() -> Self {
        unsafe { std::mem::zeroed() }
    }
}
pub unsafe fn foo_0() -> i32 {
    let mut u: Number = <Number>::default();
    u.i = 42;
    return u.i;
}
}

and the refcount model produces

#![allow(unused)]
fn main() {
pub struct Number {
    __bytes: Value<Box<[u8]>>,
}
impl Number {
    pub fn i(&self) -> Ptr<i32> {
        (self.__bytes.as_pointer() as Ptr<u8>).reinterpret_cast()
    }
    pub fn f(&self) -> Ptr<f32> {
        (self.__bytes.as_pointer() as Ptr<u8>).reinterpret_cast()
    }
}
impl Default for Number {
    fn default() -> Self {
        Number {
            // 4 is sizeof(union Number): the size of its largest member
            __bytes: Rc::new(RefCell::new(Box::from([0u8; 4]))),
        }
    }
}
pub fn foo_0() -> i32 {
    let u: Value<Number> = Rc::new(RefCell::new(<Number>::default()));
    (*u.borrow_mut()).i().write(42);
    return (*u.borrow()).i().read();
}
}

(Clone and ByteRepr impls omitted.)

Unsafe model

A union is a Rust union with #[repr(C)] and #[derive(Copy, Clone)], one field per member with the same types as a struct would have. Rust cannot derive Default for a union, so a hand-written impl zeroes the bytes, which is also what C’s zero-initialization gives. Members are read and written like struct fields; the whole function is unsafe, so no extra block is needed.

Refcount model

Rust unions are unusable from safe code, so the refcount model stores the union as one byte buffer, __bytes: Value<Box<[u8]>>, sized to the largest member, and emits one accessor method per member. Each accessor takes a pointer to the buffer and reinterprets it as the member type, returning a Ptr<T> (a Ptr to the element type for array members). Member access u.i therefore becomes a call, u.i(), and the read or write goes through the pointer: .read() and .write(v) for scalars, .upgrade().deref() for struct members whose fields are then accessed as usual.

Because every member views the same bytes, writing through one member and reading through another has C’s semantics: the bytes are reinterpreted, not converted. This is also why the type must implement ByteRepr; the accessor’s reinterpret_cast needs the member types to have a byte-level representation.

Default fills the buffer with zeros, Clone copies the buffer into a fresh Value, and ByteRepr copies the buffer in and out. The Rust struct has no pub fields, so translated code can reach the storage only through the accessors.

Warning

Accessors are broken on a reinterpreted union. Reading a Ptr<U> obtained from reinterpret_cast builds a temporary U from the bytes with from_bytes, so p.upgrade().deref().i() returns a pointer into that temporary’s buffer, which dangles as soon as the statement ends, and a write through it would never reach the original allocation (#311).

Bit-fields

Bit-fields are not implemented. VisitFieldDecl ignores the declared width, so a field such as unsigned flags : 3; is emitted as a plain u32 field: the struct compiles, but its layout, size, and the wrap-around of the field’s value differ from C.

Warning

Code that relies on bit-field layout or width (packing several fields into one word, sizeof on such a struct, storing a value wider than the field) translates silently to something else (#312).

Pointers and References

The unsafe model keeps C++ pointers as raw pointers and dereferences them directly. The refcount model replaces every pointer and reference with Ptr<T>, a weak reference plus an offset, and every dereference with a short-lived borrow of the pointee. Given

int f(int *q) {
  int b = 2;
  int *p = &b;
  *p = *q;
  return b;
}

the unsafe model produces

#![allow(unused)]
fn main() {
pub unsafe fn f_0(mut q: *mut i32) -> i32 {
    let mut b: i32 = 2;
    let mut p: *mut i32 = &mut b as *mut i32;
    *p = *q;
    return b;
}
}

and the refcount model produces

#![allow(unused)]
fn main() {
pub fn f_0(q: Ptr<i32>) -> i32 {
    let q: Value<Ptr<i32>> = Rc::new(RefCell::new(q));
    let b: Value<i32> = Rc::new(RefCell::new(2));
    let p: Value<Ptr<i32>> = Rc::new(RefCell::new(b.as_pointer()));
    p.borrow().write(q.borrow().read());
    return *b.borrow();
}
}

The rest of the page goes through the pointer operations one at a time.

Address-of

Unsafe model: &x becomes &mut x as *mut T, or &x as *const T when the pointer type is to const. Globals use &raw mut x so no reference to the static is formed. An array decays with arr.as_mut_ptr(), and the address of an element is &mut arr[i] as *mut T.

Refcount model: &x becomes x.as_pointer(), which produces a Ptr holding a weak reference to the variable’s Value. &s.field is field_ptr!(s, field), and &p->field is field_ptr!(p, field): a pointer to the field of the struct, made from the Value or the Ptr of the struct (see Pointers to fields). An array decays with arr.as_pointer() as Ptr<T>, a Ptr to element 0 of the whole array, and &arr[i] is that pointer offset by i. An array field decays with array_field_ptr!(p, arr), and p->arr[i] is read and written through that pointer offset by i.

Both models push the address down to the innermost place expression: &(cond ? x : y) becomes if cond { &mut x } else { &mut y } in the unsafe model and if cond { x.as_pointer() } else { y.as_pointer() } in the refcount one. The if yields a value copied out of whichever branch ran, not that branch’s storage, so the address has to be taken inside the branches, where the place is still known.

Dereference

Unsafe model: *p stays *p, p->x becomes (*p).x, and an assignment through a pointer is *p = v.

When the dereference is the base of an index, (*p)[i] with p: *mut Vec<i32>, Rust would have to create a &mut Vec<i32> out of the raw pointer to call Index::index, and its dangerous_implicit_autorefs lint rejects that as an error. The converter makes the reference explicit instead: EmitDeref prints (&mut (*p))[i], or (&(*p))[i] when the operator[] is const, when autoref_mut_ is set. PushExplicitAutoref sets it around the base of an overloaded subscript, which also covers a member of the pointee, (&mut (*hp)).v[i], around the range of a range-for, and around a rule placeholder marked is_index_base (see Rules IR); EmitDeref clears it once used, so nested dereferences inside the base are printed plainly.

Refcount model: a dereference cannot hand out a &T into the pointee, because nothing would bound the borrow’s lifetime, so a read copies the value out and a write copies it in. Which form is emitted follows the expression kind, what the enclosing construct expects of the dereference:

  • An rvalue use copies the value out. A scalar or pointer pointee is p.read(), and so is a whole record, *p. A field of a record pointee is copied out in a closure that borrows the record for its duration: p->x is p.with(|__s| __s.x), and p->a.b is p.with(|__s| __s.a.b); ReadField converts the record with record_ptr_ set, so that its dereference is emitted as __s. A field that is a Value of its own, or a std::unique_ptr, is copied out as well, i.e., its Rc: p->v.size() is (*p.with(|__s| __s.v.clone()).borrow()).len(). When the record is not reached through a pointer, the copy is in a block, { (*s.borrow()).x }. In all cases the record doesn’t stay borrowed for the rest of the statement, which may write to it.
  • An address-of use prints p itself.
  • An lvalue use prints nothing at once. The converter records the pointer expression as a pending dereference, and whoever consumes the lvalue, an assignment or a mapped method call, wraps it: p.write(v), or with_mut for a mutating method on a boxed pointee. This is what lets *p = v come out as a single write instead of a borrow followed by an assignment, and *p += v as { let _ptr = p.clone(); _ptr.write(_ptr.read() + v) }. A field of a record pointee is a pending dereference too, of field!(p, x), which projects the pointer to the field (see Pointers to fields): p->x = v is field!(p, x).write(v), and p->a.b = v is field!(field!(p, a), b).write(v).

with and with_mut work on any pointer, including a reinterpreted one, whose pointee is decoded from the bytes of the original allocation and, for with_mut, encoded back into them before the closure returns. A write through a reinterpreted pointer is hence visible right away through every other pointer to the same bytes.

Arithmetic and comparison

Unsafe model: p + n is p.offset(n as isize) and p - n is p.offset(-(n as isize)); p - q is (p as usize - q as usize) / ::std::mem::size_of::<T>(); ++p is p.prefix_inc() through the increment traits; p == NULL is p.is_null() and the null literal is std::ptr::null_mut(), or std::ptr::null() for a pointer to const.

Refcount model: the same operations on Ptr, p.offset(n as isize), p.clone() - q.clone() (subtraction takes its operands by value, hence the clones), p.prefix_inc(), p.is_null(), and Ptr::null(). Arithmetic only moves the offset; whether the result is in bounds is checked when it is dereferenced.

References

A C++ reference is a pointer that cannot be made to point elsewhere, and both models translate it as one. In the unsafe model a reference parameter is *mut T (*const T for const T &), an argument f(x) is f(&mut x as *mut T), and uses of the reference are *r. In the refcount model it is a Ptr<T> that is never boxed: the argument is f(x.as_pointer()), uses are r.read() and r.write(v), and returning a reference returns the Ptr (with a .clone(), since Ptr is not Copy).

Heap

new T(v) becomes Box::leak(Box::new(v)) as *mut T in the unsafe model and Ptr::alloc(v) in the refcount model; delete p becomes ::std::mem::drop(Box::from_raw(p)) and p.delete(). Array forms use a boxed slice and Ptr::alloc_array. In the refcount model the heap allocation is a leaked Rc that delete recovers, so a double delete or a delete of something that was not allocated with new panics instead of corrupting memory.

Casts

Casts are of two kinds: scalar casts, which both models spell with Rust’s as or with a small expression, and pointer casts, where the models diverge. Most casts in the input are implicit, inserted by clang, and are translated the same way as explicit ones.

Scalar casts

An integer conversion becomes expr as T (an integer literal is instead re-typed in place: 1 cast to unsigned char prints as 1_u8), and is dropped when source and target map to the same Rust type, so int to long on a platform where both are i32 prints nothing. Floating conversions are as as well. The other scalar casts have their own spellings:

  • Integer to bool: x != 0; a comparison or logical operator that already yields bool is left alone. Enum to bool compares against <E>::from(0).
  • Pointer to bool: !p.is_null().
  • Integer to enum: <E>::from(x), the From<i32> impl from the Enums page. When the operand is itself a constant of that same enum, which C++ sees as an integer being converted back to the enum, the cast is dropped and the constant is printed directly (Color::RED, not <Color>::from(Color::RED as i32)). Enum to integer is as.
  • A cast to void, used to silence an unused-variable warning, becomes a statement that only mentions the operand: &x; in the unsafe model, (*x.borrow()).clone(); in the refcount model.

Explicit static_cast, C-style, and reinterpret_cast between scalars follow the same rules; a cast to the operand’s own type is elided.

Implicit conversions to usize and isize

size_t, size_type, and ssize_t are translated as usize and isize rather than as the u64/i64 of the unsigned long/long they are typedefs of (built-in type rules, looked up on the sugared type before it is desugared). This keeps rules and output free of as usize casts on lengths and indexes, but it splits one C type in two: clang inserts no conversion between size_t and unsigned long, while usize and u64 do not mix in Rust. Given

unsigned long take_ulong(unsigned long x);

size_t sz = 20;
unsigned long r = take_ulong(sz);

the refcount model produces

#![allow(unused)]
fn main() {
let sz: Value<usize> = Rc::new(RefCell::new(20_usize));
let r: Value<u64> = Rc::new(RefCell::new(take_ulong_0(*sz.borrow() as u64)));
}

Convert(expr, implicit_convert_to) is the single place where such a cast is added: the caller passes the type the context expects, NeedsImplicitScalarCast checks that it is the same C type as the expression’s but maps to a different Rust type, and if so the expression is wrapped in (...) as <target>. Callers that pass a target are assignments and initializations (the variable’s type), call arguments (the parameter type of the callee or rule, GetParamImplicitConvertTarget, as in the example), and binary operators, which pick one Rust type for both operands (GetOperandImplicitConversionTarget).

Pointer casts

Given

uint32_t value = 0x04030201;
uint8_t *bytes = (uint8_t *)&value;
void *any = bytes;
uint8_t *back = (uint8_t *)any;

the unsafe model produces

#![allow(unused)]
fn main() {
let mut value: u32 = 67305985_u32;
let mut bytes: *mut u8 = (&mut value as *mut u32) as *mut u8;
let mut any: *mut ::libc::c_void = bytes as *mut ::libc::c_void;
let mut back: *mut u8 = any as *mut u8;
}

and the refcount model produces

#![allow(unused)]
fn main() {
let value: Value<u32> = Rc::new(RefCell::new(67305985_u32));
let bytes: Value<Ptr<u8>> =
    Rc::new(RefCell::new(value.as_pointer().reinterpret_cast::<u8>()));
let any: Value<AnyPtr> = Rc::new(RefCell::new((*bytes.borrow()).to_any()));
let back: Value<Ptr<u8>> =
    Rc::new(RefCell::new((*any.borrow()).reinterpret_cast::<u8>()));
}

Unsafe model

Every pointer cast, whether written as a C cast, static_cast, or reinterpret_cast, is a Rust as between raw pointer types. A cast that only adds or removes const changes the Rust type too, since T * is *mut T and const T * is *const T, and becomes .cast_const() or .cast_mut(). A cast that changes nothing in Rust, such as a typedef to its underlying type, is not emitted. Casts between pointers and integers are also as.

Refcount model

A Ptr<T> is a weak reference to a RefCell<T>, so it cannot simply be relabeled as a Ptr<U>: the cell it points to holds a T. A cast to another pointee type therefore produces a different kind of pointer, one that views the allocation as bytes. Three helpers from the runtime cover the cases:

  • p.reinterpret_cast::<U>() produces a Ptr<U> of the Reinterpreted kind: a byte-level view over the original allocation, with the offset counted in bytes. Reads and writes through it go through ByteRepr, which is why every record type gets a ByteRepr impl.
  • p.to_any() erases the type into an AnyPtr, the translation of void *, remembering the original type.
  • any.reinterpret_cast::<T>() recovers a Ptr<T> from an AnyPtr: the original pointer if T is the type it was erased from, a byte view otherwise (see AnyPtr casts).

Two casts do not use these helpers. An array decaying to a pointer is spelled arr.as_pointer() as Ptr<T>, where the as only names the pointer type. An upcast from a derived class to an abstract base becomes p.to_dyn::<dyn Base>(|w| w), which is the ordinary Rust unsizing coercion applied to the pointer’s weak reference (see Virtual Classes). Casts between pointers and integers use the integer cast API of Ptr.

Constness is dropped in a cast as everywhere else, so const_cast is a no-op. dynamic_cast is not supported.

Function pointers

Casting a function pointer to a different signature, which C code does to call through a generic type, wraps the function in an adapter closure that converts the arguments; see Casts on the runtime page. Storing a function pointer in a void * uses to_any() like any other pointer.

Function Pointers

Given

typedef int (*op_t)(int);

int inc(int x) { return x + 1; }

op_t pick() { return inc; }

int apply(op_t f) {
  if (f == nullptr) {
    return 0;
  }
  return f(10);
}

the unsafe model produces

#![allow(unused)]
fn main() {
pub unsafe fn inc_0(mut x: i32) -> i32 {
    return x + 1;
}
pub unsafe fn pick_1() -> Option<unsafe fn(i32) -> i32> {
    return Some(inc_0);
}
pub unsafe fn apply_2(mut f: Option<unsafe fn(i32) -> i32>) -> i32 {
    if f.is_none() {
        return 0;
    }
    return f.unwrap()(10);
}
}

and the refcount model produces

#![allow(unused)]
fn main() {
pub fn inc_0(x: i32) -> i32 {
    let x: Value<i32> = Rc::new(RefCell::new(x));
    return *x.borrow() + 1;
}
pub fn pick_1() -> FnPtr<fn(i32) -> i32> {
    return FnPtr::<fn(i32) -> i32>::new(inc_0);
}
pub fn apply_2(f: FnPtr<fn(i32) -> i32>) -> i32 {
    let f: Value<FnPtr<fn(i32) -> i32>> = Rc::new(RefCell::new(f));
    if (*f.borrow()).is_null() {
        return 0;
    }
    return (*(*f.borrow()))(10);
}
}

Unsafe model

A function pointer is Option<unsafe fn(A) -> R>: Option because it can be null, unsafe fn because every translated function is unsafe. Naming a function where a pointer is expected wraps it in Some(...), the null pointer is None, a null check is is_none(), and a call is f.unwrap()(args). Function pointers are Copy and compare with == on the function’s address.

Refcount model

A function pointer is FnPtr<fn(A) -> R>, built with FnPtr::<fn(A) -> R>::new(f). FnPtr dereferences to the function, so a call is (*f)(args); the null pointer is FnPtr::null() and the check is is_null(). FnPtr is not Copy, so storing or passing one that already lives in a variable clones it, and equality compares the address of the wrapped function, so a pointer stays equal to itself after being cast.

A lambda is also an FnPtr, so a capture-less lambda assigned to a function pointer is the lambda itself (see Lambdas).

Lambdas

A lambda becomes an FnPtr in both models. Its type is FnPtr<fn(A) -> R>, with the signature of the lambda’s call operator, so it can be written wherever C++ names the closure type: a variable, a struct field, a parameter of an instantiated template, or decltype. A call is f.call(args).

Given

int total = 0;
auto accumulate = [total](int x) mutable {
  total += x;
  return total;
};
accumulate(1);

the unsafe model produces

#![allow(unused)]
fn main() {
let mut total: i32 = 0;
let mut accumulate: FnPtr<fn(i32) -> i32> = lambda_unsafe!(
    {
        let total: i32 = total;
    },
    |x: i32| -> i32 {
        total += x;
        return total;
    }
);
(unsafe { accumulate.call(1) });
}

and the refcount model produces

#![allow(unused)]
fn main() {
let total: Value<i32> = Rc::new(RefCell::new(0));
let accumulate: Value<FnPtr<fn(i32) -> i32>> = Rc::new(RefCell::new(lambda!(
    {
        let total: Value<i32> = Rc::new(RefCell::new((*total.borrow())));
    },
    |x: i32| -> i32 {
        let x: Value<i32> = Rc::new(RefCell::new(x));
        (*total.borrow_mut()) += (*x.borrow());
        return (*total.borrow());
    }
)));
({ (*accumulate.borrow()).call(1) });
}

The lambda macros

A lambda with captures is written with lambda! in the refcount model and lambda_unsafe! in the unsafe model. Both take two arguments:

  • a block with one let per capture, giving its name, its type and the value it is initialized with when the lambda is created;
  • a closure with the lambda’s parameters, its return type and the translated body.

The macro declares a hidden struct with one field per capture, and makes the body the call method of that struct. The body is written with the names of the captures, as the C++ body is, so the macro adds self. in front of each capture name in it, which makes the name refer to the field of the struct. The lambda! above expands to

#![allow(unused)]
fn main() {
{
    struct __Lambda {
        total: Value<i32>,
    }
    impl __Lambda {
        fn call(&self, x: i32) -> i32 {
            let x: Value<i32> = Rc::new(RefCell::new(x));
            (*self.total.borrow_mut()) += (*x.borrow());
            return (*self.total.borrow());
        }
    }
    FnPtr::<fn(i32) -> i32>::from_lambda(
        __Lambda {
            total: Rc::new(RefCell::new((*total.borrow()))),
        },
        __Lambda::call,
    )
}
}

The struct is local to the expression and its type is erased by the FnPtr, so it never appears in the translated program. The initializers are not part of the method: they are evaluated where the lambda is written, so total there is the enclosing variable, while total in the body is self.total.

The two macros differ in how the method receives the struct. lambda! takes it by immutable reference, since the captures of the refcount model are Values and are written through their cells. lambda_unsafe! takes it by mutable reference, since the captures of the unsafe model are plain fields, and wraps the body in unsafe.

Every use of a capture name in the body gets the self., so the body cannot give that name to something else: a variable declared in the body, or a parameter of a closure in it, with the name of a capture is a compile error.

Captures

The captures are the ones of the C++ closure type, which clang computes also for [=] and [&].

CaptureUnsafe modelRefcount model
[x]TValue<T>
[&x]*mut TPtr<T>
[n = expr]TValue<T>
[this]*mut SValue<Ptr<S>>
[*this]SValue<S>

A capture by copy holds its own copy, made when the lambda is created, and keeps its value between calls. A capture by reference holds a pointer to the variable and is dereferenced in the body:

int base = 10;
auto add_base = [&base](int x) { return x + base; };
#![allow(unused)]
fn main() {
let mut base: i32 = 10;
let mut add_base: FnPtr<fn(i32) -> i32> = lambda_unsafe!(
    {
        let base: *mut i32 = &mut base;
    },
    |x: i32| -> i32 {
        return ((x) + (*base));
    }
);
}

A captured this is the capture this_. Inside the body this refers to it, and members are accessed through that pointer.

A constant that the body uses without capturing, such as a const int local read by value, is replaced by its initializer.

Lambdas without captures

A lambda without captures needs no struct and is a function pointer built from a closure:

auto one = [](int x) { return x + 1; };
#![allow(unused)]
fn main() {
let mut one: FnPtr<fn(i32) -> i32> = FnPtr::<fn(i32) -> i32>::new(|x: i32| -> i32 {
    unsafe {
        return ((x) + (1));
    }
});
}

Converting it to a function pointer gives, in the refcount model, the lambda itself, which already has the type of the function pointer. The unsafe model translates function pointers as Option<unsafe fn(A) -> R>, so the conversion emits Some(|...| ...) with the closure again (see Function Pointers). Lambdas with captures cannot be converted to function pointers, as in C++.

A lambda without captures can be default-constructed since C++20. The translation of decltype(one) other; emits the closure again, as do the fields of that type in the Default of a struct.

Limitations

  • Copying a lambda does not copy its captures: the copy shares them with the original, so a mutable lambda and its copy update the same state, and the copy constructors of the captures do not run.
  • Moving a lambda does not run the move constructors of its captures.
  • Destroying a lambda does not run the destructors of its captures.
  • A lambda that is default-constructed by a type translated with a rule is wrong. The rule only has the type of the lambda, FnPtr<fn(A) -> R>, whose default is the null pointer and not the lambda, so calling it panics. This affects, for example, std::set<int, decltype(cmp)> s;, where the set constructs the comparator.
  • Generic lambdas, whose call operator is a template, are not translated.
  • A default-constructed lambda whose body declares a static local is rejected, as each emitted closure would get its own copy of the variable.
  • The parameter and return types of a lambda must implement FnPtrArg, like those of any FnPtr.

Special-cased Library Types

Library types are translated by type rules, and for most of them the converter does nothing beyond applying the rule. A few types also have code of their own in the converter, which decides how their values are dereferenced, iterated, or initialized. Some of it could be moved into rules. This page lists what the converter does today; the rest of the type comes from its rule module under rules/.

std::unique_ptr

The type rule maps std::unique_ptr<T> to Option<Box<T>> in the unsafe model and Option<Value<T>> in the refcount model, and std::make_unique and std::move are rules too. What the converter special-cases (IsUniquePtr in converter_lib) is everything that treats a unique_ptr as a pointer:

Given

std::unique_ptr<int> x1 = std::make_unique<int>(0);
std::unique_ptr<int> x2 = std::make_unique<int>(0);
*x2 = 1;
x1 = std::move(x2);
int *raw = &*x1;

the unsafe model produces

#![allow(unused)]
fn main() {
let mut x1: Option<Box<i32>> = Some(Box::new(0));
let mut x2: Option<Box<i32>> = Some(Box::new(0));
*x2.as_deref_mut().unwrap() = 1;
x1 = x2;
let mut raw: *mut i32 = &mut (*x1.as_deref_mut().unwrap()) as *mut i32;
}

and the refcount model produces

#![allow(unused)]
fn main() {
let x1: Value<Option<Value<i32>>> =
    Rc::new(RefCell::new(Some(Rc::new(RefCell::new(0)))));
let x2: Value<Option<Value<i32>>> =
    Rc::new(RefCell::new(Some(Rc::new(RefCell::new(0)))));
*(*x2.borrow_mut()).as_ref().unwrap().borrow_mut() = 1;
*x1.borrow_mut() = (*x2.borrow_mut()).take();
let raw: Value<Ptr<i32>> =
    Rc::new(RefCell::new((*x1.borrow()).as_pointer()));
}

*p and p->x are overloaded operator calls in C++; the converter emits the as_deref_mut().unwrap() and as_ref().unwrap().borrow_mut() forms instead of an operator call, &*p becomes the raw pointer or as_pointer(), and std::move of a unique_ptr is a plain move or a take(). In the unsafe model p == nullptr becomes p.is_none(); the refcount model does not special-case it and emits is_null() as for any pointer, which no test exercises on an Option<Value<T>>. A struct with a unique_ptr field does not derive Copy.

Iterators

Iterator types come from rules (std::vector<T>::iterator maps to *mut T or Ptr<T>, std::map<K, V>::iterator to UnsafeMapIterator<K, V> or RefcountMapIter<K, V>), and the converter classifies them by GetStrongestIteratorCategory: rule types marked as refcount pointers are contiguous iterators and are handled exactly like a Ptr<T>, and the map iterator types are bidirectional. The classification drives a few decisions.

Given

std::map<int, double> m;
double sum = 0;
for (const auto &i : m) {
  sum += i.second;
}
auto it = m.begin();
sum += it->second;

the unsafe model produces

#![allow(unused)]
fn main() {
for i in UnsafeMapIterator::begin(&m as *const BTreeMap<i32, Box<f64>>) {
    sum += *i.second();
}
let mut it: UnsafeMapIterator<i32, f64> =
    UnsafeMapIterator::begin(&m as *const BTreeMap<i32, Box<f64>>);
sum += *it.second();
}

and the refcount model produces

#![allow(unused)]
fn main() {
for i in RefcountMapIter::begin(m.as_pointer()) {
    *sum.borrow_mut() += *i.second().borrow();
}
let it: Value<RefcountMapIter<i32, f64>> =
    Rc::new(RefCell::new(RefcountMapIter::begin(m.as_pointer())));
*sum.borrow_mut() += *(*it.borrow()).second().borrow();
}

it->second on a bidirectional iterator (map iterator) is not a pointer dereference plus a field access, since the map iterator types have no pointer to hand out; the converter emits the iterator itself and the field rule turns the access into an accessor call.

The loop variable of a range-for over a std::map is the map iterator itself, an entry with first()/second() accessors, not a pointer to an element; the converter remembers such variables in map_iter_decls_ so that uses of them are not dereferenced.

A converting-constructor call that only wraps an iterator does not clone it (PushSuppressIteratorClone). libstdc++ and libc++ differ in whether such a wrapping constructor appears in the AST, so skipping the clone keeps the output identical on Linux and macOS.

IsIteratorType recognizes any record that declares an iterator_category typedef.

std::array

std::array<T, N> maps to Vec<T> (see rules/array), so an initializer {1, 2, 3} becomes vec![1, 2, 3]. The converter knows the type by name in three places: the default value of an uninitialized std::array variable is built element by element from N, a struct with a std::array field does not derive Default, and it does not derive Copy either since the field is a Vec.

Warning

An empty initializer, std::array<int, 3> a = {};, becomes vec![], a vector of length 0, where C++ value-initializes N elements; indexing it panics (#313).

std::string and streams

std::string maps to Vec<libc::c_char> in the unsafe model and Vec<u8> in the refcount model through its rules; the converter itself only special-cases string literals (their type in an initializer, and ASCII escaping) and range-for over a string. std::ostream calls (std::cout << x) are detected with IsCallToOstream and translated by a dedicated path rather than by rules; that path is described with printf under Expressions.