Introduction
Cpp2Rust translates C++ to fully safe Rust automatically. It is a syntax-driven translator based on clang’s AST.
Cpp2Rust’s algorithm is described in the paper Cpp2Rust: Automatic Translation of C++ to Safe Rust published at PLDI 2026.
Overview
Cpp2Rust first parses the input C++ file(s) with clang and produces an AST. It
then traverses the AST and emits Rust code as strings, inserting calls to the
libcc2rs runtime library where needed (e.g., for raw pointer semantics).
Finally, the Rust code is pretty-printed using rustfmt to a single .rs file.
By default the reference counting model is used, which produces fully safe
Rust. A generator of unsafe Rust is also available through the --model=unsafe
command line argument for debugging and performance comparisons.
Runtime library (libcc2rs)
The generated code relies on a runtime library
designed to simplify the translation process. C pointers are converted into the
Ptr<T> type provided by libcc2rs. Ptr<T> models C pointer semantics,
including null, arithmetic, and aliasing, while satisfying Rust’s borrow checker
through checked run-time operations.
Building
Requirements
On Ubuntu, install the required dependencies with:
sudo apt install libclang-22-dev clang++-22 ninja-build cmake
pip install ruff==0.15.22
Build
mkdir build
cd build
cmake -GNinja ..
ninja
ninja check
Usage
Translate a single file
./build/cpp2rust/cpp2rust --file=<file>.cpp -o=<file>.rs
By default, the reference counting model is used (fully safe output). To generate unsafe Rust instead:
./build/cpp2rust/cpp2rust --file=<file>.cpp -o=<file>.rs --model=unsafe
Minimal example. Given hello.cpp:
#include <cstdio>
int main() {
printf("hello world\n");
return 0;
}
Running ./build/cpp2rust/cpp2rust --file=hello.cpp -o=hello.rs produces:
pub fn main() {
std::process::exit(main_0());
}
fn main_0() -> i32 {
println!("hello world");
return 0;
}
Compile and run with:
rustc hello.rs -L build/libcc2rs-target/release
./hello
Translate a whole program
First generate a
compile_commands.json
for your project. With CMake this is one extra flag:
cmake -DCMAKE_EXPORT_COMPILE_COMMANDS=ON ..
Then run:
./build/cpp2rust/cpp2rust --dir=<dir> -o <output>.rs
<dir> must be the directory that contains compile_commands.json.
Test Suite
# Run all tests
ninja check
# Run only the unit tests
ninja check-unit
# Run libcc2rs unit tests
ninja check-libcc2rs
# Run libcc2rs-macros unit tests
ninja check-libcc2rs-macros
# Regenerate expected output for unit tests after intentional changes
REPLACE_EXPECTED=1 ninja check-unit
Overview
Translation rules describe how C++ library APIs are mapped to Rust. Each rule
module lives in the rules/ directory and pairs a C++ source file (src.cpp)
with its Rust translation for each model (tgt_unsafe.rs and
tgt_refcount.rs).
Every rule is expressed as ordinary, compilable C++ and Rust source code, free
of anything platform dependent. Both sides are run through real compilers at
build time, so a rule that does not compile fails the build, and the
platform-specific spellings (for example bool canonicalizing to _Bool) are
derived by the compiler on the host rather than written by hand. The same rule
sources work on every platform cpp2rust builds on.
Rules go through a build-time compilation pipeline before cpp2rust can use
them:
- You author a rule module: C++ patterns in
src.cppand Rust targets intgt_unsafe.rs/tgt_refcount.rs. - At build time, two preprocessors compile the module into Rules IR under
<build>/rules/<module>/:cpp-rule-preprocessorcompiles the C++ side intoir_src.json, andrule-preprocessorcompiles the Rust side intoir_unsafe.jsonandir_refcount.json. - At startup,
cpp2rustloads the Rules IR files and indexes the rules by the canonical signature of the C++ construct they match.
The rest of this part covers each stage:
- Rule Format: the files that make up a rule module and how the two models are layered.
- Writing Rules: how to write rules for functions, methods, operators, types, constants, and variadics.
- Compat Shims: how macro-based libc APIs like
errnoandFD_SETare rewritten into matchable function calls. - Conventions: naming and style conventions rule authors must follow.
- The Rule Preprocessors: the two build-time tools that compile rules to the Rules IR.
- The Rules IR: the JSON format the preprocessors emit.
- Loading and Matching: how
cpp2rustloads the Rules IR and matches rules against the input AST. - The Matching Engine: how a candidate rule’s signature is unified against the input.
- Rule Rewriting: how rule bodies are adapted at application
time, in particular the
with_mutrewrite.
Rule Format
A rule module is a directory under rules/, usually named after the header or
library it covers (rules/unistd, rules/vector, rules/string, …). It
contains:
src.cppand/orsrc.c: the C++ (or C) side of each rule.tgt_unsafe.rs: the Rust targets for the unsafe model.tgt_refcount.rs: the Rust targets for the reference counting model (optional, see below).
A rule is a pair of same-named functions on the two sides. Names determine the rule kind:
f1,f2, … are expression rules: they map a C++ call, member access, constructor, or constant to a Rust expression.t1,t2, … are type rules: they map a C++ type to a Rust type.
Expression rules
On the C++ side, an fN function must have a body that is exactly one return
statement. The returned expression is the pattern: the preprocessor resolves the
callee of that expression (the function, method, constructor, enum constant,
or macro being used) and that becomes the rule’s matching key. The function
parameters stand for the arguments at the call site.
// rules/unistd/src.cpp
int f4(const char *pathname) { return unlink(pathname); }
On the Rust side, the same-named function gives the replacement expression.
Parameters must be named a0, a1, … and correspond positionally to the C++
parameters:
#![allow(unused)]
fn main() {
// rules/unistd/tgt_unsafe.rs
unsafe fn f4(a0: *const libc::c_char) -> i32 {
libc::unlink(a0)
}
}
#![allow(unused)]
fn main() {
// rules/unistd/tgt_refcount.rs
fn f4(a0: Ptr<u8>) -> i32 {
match nix::unistd::unlink(a0.to_rust_string().as_str()) {
Ok(()) => 0,
Err(__e) => {
libcc2rs::cpp2rust_errno().write(__e as i32);
-1
}
}
}
}
When the converter encounters unlink(x) in the input, it emits the rule body
with the translated x substituted for a0.
Type rules
On the C++ side, a tN rule is a type alias (using or typedef). On the Rust
side, it is a zero-argument function whose return type is the mapped Rust type
and whose body is the default initializer for that type:
// rules/vector/src.cpp
template <typename T1> using t1 = std::vector<T1>;
#![allow(unused)]
fn main() {
// rules/vector/tgt_unsafe.rs
fn t1<T1>() -> Vec<T1> {
Vec::new()
}
}
Model layering
The loader always reads ir_unsafe.json first. When translating with the
reference counting model, it then overlays ir_refcount.json on top: entries
with the same rule name replace the unsafe ones.
This means tgt_refcount.rs only needs to contain the rules that differ from
the unsafe model. For example, the __builtin_mul_overflow rule in
rules/builtin has a pointer out-parameter (a2 below), so the two models
translate it differently: in the unsafe model a2 is a raw *mut i64 written
through a deref, while in the refcount model it is a Ptr<i64> written through
Ptr::write. The other two arguments are identical in both models:
#![allow(unused)]
fn main() {
// rules/builtin/tgt_unsafe.rs
unsafe fn f9(a0: i64, a1: i64, a2: *mut i64) -> bool {
let (val, ovf) = a0.overflowing_mul(a1);
*a2 = val;
ovf
}
}
#![allow(unused)]
fn main() {
// rules/builtin/tgt_refcount.rs
fn f9(a0: i64, a1: i64, a2: Ptr<i64>) -> bool {
let (val, ovf) = a0.overflowing_mul(a1);
a2.write(val);
ovf
}
}
The module’s other rules (byte swaps, __builtin_expect, …) translate
identically in both models, so they appear only in tgt_unsafe.rs and the
refcount model inherits them. A module where no rule needs a refcount-specific
translation can omit tgt_refcount.rs entirely.
C and C++ sources
A module may have both src.c and src.cpp; both are preprocessed and merged
into one ir_src.json. Defining the same rule name in both files is a hard
error, so numbering must not collide.
This split is necessary because rules match on the exact canonical signature of
the callee, and some libc functions have different signatures in C and C++.
For example, C has a single char *strchr(const char *, int), while C++
replaces it with const-correct overloads such as
const char *strchr(const char *, int). Since the signatures differ,
rules/cstring defines one rule per language:
// rules/cstring/src.cpp
const char *f6(const char *a0, int a1) { return strchr(a0, a1); }
// rules/cstring/src.c
char *f5(const char *a0, int a1) { return (strchr)(a0, a1); }
The C++ rule matches strchr calls in code translated as C++, the C rule
matches them in code translated as C.
rules/builtin uses this to cover both languages: src.cpp defines f9/f10
for the C++ __builtin_mul_overflow (returning bool) while src.c defines
f12/f13 for the C version (returning int); their Rust bodies are
identical.
The rules crate
The whole rules/ tree is a single Rust crate. rules/build.rs walks the tree,
collects every tgt_*.rs, and generates rules/src/modules.rs with one
#[path = ...] module per file. Building the crate therefore type-checks every
rule body against the crates the rule targets call into, which are declared as
dependencies in rules/Cargo.toml (libcc2rs, libc, nix, …). The Rust
rule preprocessor compiles exactly this crate to resolve types in rule bodies.
rules/src/ is the only subdirectory that is not a rule module.
Writing Rules
This page shows how to write rules for each kind of C++ construct. In every case
the recipe is the same: write an fN (or tN) function on the C++ side whose
single return statement exercises the construct, and a same-named function on
the Rust side giving the translation.
Free functions
// rules/stat/src.cpp
int f1(const char *pathname, struct stat *statbuf) {
return stat(pathname, statbuf);
}
#![allow(unused)]
fn main() {
// rules/stat/tgt_refcount.rs
fn f1(a0: Ptr<u8>, a1: Ptr<Stat>) -> i32 {
match nix::sys::stat::stat(a0.to_rust_string().as_str()) {
Ok(__s) => {
a1.with_mut(|__st| *__st = Stat::from_libc(&__s));
0
}
Err(__e) => {
libcc2rs::cpp2rust_errno().write(__e as i32);
-1
}
}
}
}
Rule bodies may be arbitrarily complex; multi-statement bodies are wrapped in a block when spliced into the output.
return statements are prohibited in Rust rule bodies (the preprocessor rejects
them); produce the result as a tail expression instead. The body is not emitted
as a function of its own: it is spliced inline into the generated code as a
block expression, so a return would not end the rule, it would return from
whatever generated function the rule happens to be expanded in.
When the pattern’s type cannot be named, the rule uses an auto return type:
rules/iomanip writes auto f1(int n) { return std::setw(n); } because
std::setw returns an unspecified type.
Methods
There is no special syntax for member functions: write a free function that
takes the receiver as its first parameter and calls the method on it. On the
Rust side the receiver is a0.
// rules/vector/src.cpp
template <typename T1> std::size_t f2(const std::vector<T1> &o) {
return o.size();
}
#![allow(unused)]
fn main() {
// rules/vector/tgt_unsafe.rs
unsafe fn f2<T1>(a0: Vec<T1>) -> usize {
a0.len()
}
}
Template rules use generic parameters named T1, T2, … on both sides,
matched positionally. The rule is written against the open template
std::vector<T1>, with T1 left as a placeholder, so a single rule covers
every instantiation: when the input program calls size() on, say, a
std::vector<int>, the matcher binds T1 = int.
Static member functions
A static member function is also written with a receiver parameter, which exists
only to name the class. The call site has no receiver argument, so the Rust side
drops it and numbers the remaining parameters from a0; here there are none:
// rules/limits/src.cpp
template <typename T1> T1 f1(std::numeric_limits<T1> &a0) { return a0.max(); }
#![allow(unused)]
fn main() {
// rules/limits/tgt_unsafe.rs
unsafe fn f1<T1: HasMinMax>() -> T1 {
<T1>::MAX
}
}
(HasMinMax is a helper trait defined alongside the rules in the same file.)
Constructors
Constructors are functions returning the type by value, one rule per overload:
// rules/string/src.cpp
std::string f7(const char *s, std::size_t n) { return std::string(s, n); }
std::string f9(std::size_t n, char ch) { return std::string(n, ch); }
Overloads that differ in value category are distinct rules too: rules/vector
has separate rules for push_back(const T1 &) and push_back(T1 &&).
No destructor rules exist so far: the STL and libc APIs covered by the current
rules have not needed any, since their types map to Rust types whose Drop
implementations already do the right thing.
Operators
Write operators with explicit operator call syntax, in member form
(x.operator@(...)) or free form (operator@(a, b)):
// rules/map/src.cpp
template <typename T1, typename T2>
T2 &f1(std::map<T1, T2> &o, const T1 &key) { return o.operator[](key); }
template <typename T1, typename T2>
bool f11(typename std::map<T1, T2>::iterator a,
typename std::map<T1, T2>::iterator b) {
return operator!=(a, b);
}
Post-increment is distinguished from pre-increment by the usual dummy int
parameter: a0.operator++(a1) versus it.operator++(). Conversion operators
use the same explicit syntax: a0.operator T1 &() in rules/functional matches
the conversion of a std::reference_wrapper<T1> back to a reference. Field
accesses are rules of their own, matched by the field: it->first and
it->second through iterators, plain o.second on a pair (rules/map,
rules/pair).
Callable arguments
A rule parameter may be a callable. Function pointers are spelled directly; for
a lambda, whose type cannot be written, the rule declares a file-scope lambda
and takes decltype(lambda):
// rules/algorithm/src.cpp
auto lambda = [](const T2 &a, const T2 &b) { return false; };
void f6(T1 first, T1 last, decltype(lambda) comp) {
return std::stable_sort(first, last, comp);
}
#![allow(unused)]
fn main() {
// rules/algorithm/tgt_unsafe.rs
unsafe fn f6<T1: Ord, T2>(a0: *mut T1, a1: *mut T1, a2: &mut T2)
where
T2: FnMut(&T1, &T1) -> bool,
{ ... }
}
T1 and T2 are not template parameters here but file-scope helper structs
modelling an iterator and its value type; being named like generics, they bind
as T1/T2 at the use site. The function pointer version of the comparator is
a separate rule (f7).
Iterators
There is no iterator abstraction: an iterator type gets a type rule, and every
operation on it its own expression rule (operator*, operator++,
operator!=, …). What the type maps to is up to the rule:
std::string::iterator becomes a plain pointer (*mut libc::c_char unsafe,
Ptr<u8> refcount), while std::map iterators become the runtime types
libcc2rs::UnsafeMapIterator/MapIterator. Dependent iterator types are named
with typename:
// rules/map/src.cpp
template <typename T1, typename T2>
using t2 = typename std::map<T1, T2>::const_iterator;
Types
A type rule has two halves. On the C++ side, declare a type alias named tN for
the C++ type being mapped. On the Rust side, write a function with the same name
that takes no arguments: its return type is the Rust type that the C++ type
maps to, and its body is the default value the generated code uses when it
needs to construct one (e.g. for an uninitialized variable). Reference and
pointer variants of a type each get their own rule:
// rules/iostream/src.cpp
using t1 = std::ostream;
using t2 = std::ostream &;
using t3 = std::ostream *;
C structs use typedef instead of using:
// rules/stat/src.cpp
typedef struct stat t1;
#![allow(unused)]
fn main() {
// rules/stat/tgt_unsafe.rs
fn t1() -> ::libc::stat { unsafe { std::mem::zeroed() } }
}
#![allow(unused)]
fn main() {
// rules/stat/tgt_refcount.rs
fn t1() -> libcc2rs::Stat { Default::default() }
}
A type rule may map to the sentinel type libcc2rs::IgnoreRule, meaning “this
model has no special mapping for the type”; the converter then falls back to its
normal type conversion. This is useful when only one model needs a custom
mapping: rules/carray maps multi-dimensional C arrays to nested boxed slices
in the refcount model, while its tgt_unsafe.rs targets are IgnoreRule so the
unsafe model keeps the default array conversion.
Enum values, constants, and macros
Constants are fN functions that take no arguments and return the constant, one
rule per value:
// rules/fcntl/src.cpp
int f3(void) { return O_CREAT; }
int f4(void) { return O_TRUNC; }
#![allow(unused)]
fn main() {
// rules/fcntl/tgt_unsafe.rs
unsafe fn f3() -> i32 { ::libc::O_CREAT }
}
For macros that expand to integer literals, the preprocessor records the macro
name rather than the value, so O_CREAT in the input matches this rule by
name. Enum constants and global variables (e.g. std::cout) are matched by
their qualified name. A global and its address are separate rules:
rules/iostream maps both std::cout (f1) and &std::cout (f3).
A variable of class type, such as std::cout or std::strong_ordering::less,
must be returned by reference. Returned by value, the pattern is a copy
construction, so the rule keys on the copy constructor instead of the variable
and matches every copy of that type:
// rules/compare/src.cpp
const std::strong_ordering &f1() { return std::strong_ordering::less; }
const std::strong_ordering &f2() { return std::strong_ordering::equal; }
The Rust side still returns the value; the reference only exists to keep the C++ pattern free of the copy.
Integer-literal macros are the only macros matchable directly. Macros whose
expansions are platform internals with no stable callee, such as errno or
FD_SET, are first rewritten into calls to synthetic cpp2rust_* functions by
the compat shims; rules then match the shim call.
Variadic functions
The C++ side uses a template parameter pack rather than a C-style ...
parameter, out of necessity: a function that takes ... cannot forward its
variadic arguments to another call, so a rule like
int f1(int a0, int a1, ...) { return fcntl(a0, a1, ...); }
is not expressible. A parameter pack can be forwarded (args...), which is
exactly what the rule body needs to do. The Rust side takes a trailing parameter
that must be typed &[VaArg] and named va:
// rules/fcntl/src.cpp
template <typename... Args>
int f1(int a0, int a1, Args... args) {
return fcntl(a0, a1, args...);
}
#![allow(unused)]
fn main() {
// rules/fcntl/tgt_refcount.rs
fn f1(a0: i32, a1: i32, va: &[VaArg]) -> i32 { ... }
}
Bodies read the arguments through the va-args API in libcc2rs (VaArg,
VaList, the VaArgGet accessors, format_c).
Constructing from forwarded arguments
Functions like emplace_back forward their arguments to a constructor. Rust has
no equivalent, so the rule receives the finished value instead: the Rust side
takes a trailing parameter named init, and the converter builds it at the call
site from the arguments after the fixed ones. The C++ side spells the pack as
Init<T, Args>, a transparent alias
(template <typename T, typename A> using Init = A;) whose T names the type
to build:
// rules/vector/src.cpp
template <typename T1, typename... Args>
T1 &f112(std::vector<T1> &o, Init<T1, Args> &&...args) {
return o.emplace_back(std::forward<Args>(args)...);
}
#![allow(unused)]
fn main() {
// rules/vector/tgt_unsafe.rs
unsafe fn f112<T1>(a0: &mut Vec<T1>, init: T1) {
let __init = init;
a0.push(__init)
}
}
T must be one of the callee’s template arguments. The preprocessor records its
position as a (depth, index) pair, and at a call like v.emplace_back(4, 5) on
a std::vector<Point> the converter reads the template argument at that
position from the resolved callee (Point). It then asks Sema which constructor
builds a Point from (4, 5) and substitutes the converted construction for
init. Binding init to a local before touching a0 keeps the construction
from overlapping a borrow of the container.
Passthrough rules
When a call should be forwarded verbatim to the same-named function in Rust’s
libc crate, the Rust target can be an extern declaration instead of a body:
// rules/fcntl/src.cpp
template <typename... Args>
int f1(int a0, int a1, Args... args) {
return fcntl(a0, a1, args...);
}
#![allow(unused)]
fn main() {
// rules/fcntl/tgt_unsafe.rs
unsafe extern "C" {
fn f1(a0: i32, a1: i32, ...) -> i32;
}
}
The converter then emits a direct libc::fcntl(...) call at the call site.
Platform-specific rules
Gate the C++ side with the usual preprocessor conditionals and the Rust side
with #[cfg(...)]; the two must agree so that the rule name sets line up:
// rules/socket/src.c
#ifdef __linux__
int f4(void) { return SOCK_CLOEXEC; }
#endif
#![allow(unused)]
fn main() {
// rules/socket/tgt_unsafe.rs
#[cfg(target_os = "linux")]
unsafe fn f4() -> i32 {
libc::SOCK_CLOEXEC
}
}
The Rust preprocessor evaluates #[cfg] attributes against the host target
(only target_os = linux|macos and target_arch = x86_64|x86 are accepted) and
drops non-matching rules.
Mutually exclusive platform branches use #elif with disjoint rule numbers:
rules/errno defines f91 to f135 under __linux__ and f136 to f153
under __APPLE__. Feature-test macros a pattern needs must come before the
includes, as with #define _GNU_SOURCE in rules/socket/src.c.
Pattern resolution limits
The preprocessor resolves a template pattern by instantiating its template parameters with synthesized types (The Rule Preprocessors):
- A bare
T1becomes an empty struct, so the pattern cannot use members, operators, or nested types ofT1. - A parameter pack instantiates to the empty pack.
- A non-type parameter is pinned to the value
1.
Unqualified callees are looked up in namespace std first and in the global
scope only when std has no match, so an unqualified name that exists in both
resolves to the std one.
Compat Shims
Rule matching needs a resolvable callee. The preprocessor keys every expression
rule on the function, method, constructor, constant, or global that the
pattern’s return expression resolves to, and the only macros it can record are
those that
expand to an integer literal,
which match by macro name. Any other macro is invisible to the rule system: by
the time clang has built the AST, the macro is gone and only its expansion
remains.
That is a problem for a small set of libc APIs that are specified as macros over platform internals:
errnois an object-like macro; glibc expands it to(*__errno_location()), macOS to(*__error()).assertexpands to a conditional that stringifies the condition and calls a platform-specific failure handler with file and line arguments.FD_SET,FD_CLR,FD_ISSET, andFD_ZEROexpand to bit manipulation on thefd_setrepresentation, through helpers that differ per platform.ntohl,ntohs,htonl, andhtonsexpand to byte-swap builtins or to nothing at all, depending on endianness.
There is no stable, platform-independent callee here to key a rule on. The
compat headers in cpp2rust/compat/ fix this by rewriting each such macro
into a call to a synthetic, well-known function before matching happens.
How the shims work
cpp2rust/compat is injected as a system include directory ahead of the
platform headers in every clang invocation the project makes: both when
cpp2rust parses the input program and when cpp-rule-preprocessor compiles
rule sources. The shared flag list lives in cpp2rust/compat/platform_flags.h
(getPlatformClangBeginFlags), and the directory path is baked in at build time
via the COMPAT_INCLUDE_DIR definition.
A shim header sits at the same relative path as the real header it shadows
(errno.h, sys/select.h, arpa/inet.h, …), so an ordinary
#include <errno.h> finds the shim first. The header then:
- pulls in the real platform header with
#include_next(a GNU extension; the shared flags pass-Wno-gnu-include-nextfor it), #undefs the macro,- declares a
cpp2rust_*shim function, - redefines the macro to call the shim.
cpp2rust/compat/errno.h in full:
#include_next <errno.h>
#undef errno
int *cpp2rust_errno(void);
#define errno (*cpp2rust_errno())
The redefinition keeps errno an lvalue by dereferencing the returned pointer,
so both reads and assignments like errno = 0 still parse; what the matcher
sees in either case is a call to int *cpp2rust_errno().
Because the input program and the rule sources are compiled with the same shim headers, both sides canonicalize to the same signature, and an ordinary expression rule matches it:
// rules/errno/src.c
#include <errno.h>
int *f1(void) { return cpp2rust_errno(); }
The shim functions are declared but never defined on the C side. They only exist so that the callee resolves; translation replaces the call with the rule body, so no C implementation is ever linked. Whatever the shim is supposed to do is supplied by the Rust targets:
#![allow(unused)]
fn main() {
// rules/errno/tgt_unsafe.rs
unsafe fn f1() -> *mut i32 {
libcc2rs::cpp2rust_errno_unsafe()
}
}
#![allow(unused)]
fn main() {
// rules/errno/tgt_refcount.rs
fn f1() -> Ptr<i32> {
libcc2rs::cpp2rust_errno()
}
}
In the unsafe model libcc2rs::cpp2rust_errno_unsafe wraps the real platform
errno location (__errno_location on Linux, __error on macOS). The refcount
model instead virtualizes errno as a thread-local Value<i32> inside
libcc2rs. Nothing else writes that cell, so it is a discipline of the rules:
every refcount rule that translates a call that can fail must write the error
code into it on the failure path, as
libcc2rs::cpp2rust_errno().write(__e as i32) in the
stat rule; a rule that skips the write breaks programs
that check errno.
A rule pattern may spell either the macro or the shim directly; the two are
identical after expansion. rules/errno and rules/assert call the shim by
name, while rules/arpa_inet and rules/select write the macro form:
// rules/select/src.cpp
void f2(int fd, fd_set *set) { return FD_SET(fd, set); }
The current shims
| Header | Macros | Shim functions | Rules |
|---|---|---|---|
assert.h | assert | cpp2rust_assert_fail(bool) | rules/assert maps it to assert!(a0) |
errno.h | errno | cpp2rust_errno() | rules/errno, see above |
arpa/inet.h | ntohl, ntohs, htonl, htons | cpp2rust_ntohl(x), … | rules/arpa_inet maps them to u32::from_be, u16::to_be, … |
sys/select.h | FD_SET, FD_CLR, FD_ISSET, FD_ZERO | cpp2rust_fd_set(fd, set), … | rules/select maps them to libc::FD_SET(...) (unsafe) or CFdSet methods (refcount) |
Note how the shim also normalizes the shape of the API. C’s assert is a
macro precisely so it can stringify its condition and capture file and line; the
shim reduces it to a plain void(bool) function, and the Rust side regains the
diagnostics by mapping it to the assert! macro.
Adding a new shim
To make another macro-based API matchable:
- Create the header in
cpp2rust/compat/at the same relative path as the platform header that defines the macro. - Follow the pattern above:
#include_nextthe real header,#undefthe macro, declare acpp2rust_<name>function with the macro’s effective signature, and redefine the macro to call it. - Write rules for the shim in a
rules/module as for any other function, including the corresponding header insrc.c/src.cpp. - If a model needs runtime support (as refcount errno does), implement it in
libcc2rsand call it from the rule target.
Keep the shim’s signature platform-independent; the whole point is that both sides of every rule see one canonical declaration on every platform.
Related normalization
The same shared flag list also passes -D_FORTIFY_SOURCE=0, which keeps glibc
from substituting fortified variants (__printf_chk and friends) for standard
calls. Like the shims, this ensures that calls in the input program resolve to
the standard declarations the rules are written against.
Conventions
Most of these conventions are enforced by the preprocessors, and violating them fails the build; the notes below call out the ones that are not checked.
Naming
| Element | C++ side | Rust side |
|---|---|---|
| Expression rule | f1, f2, … | same name |
| Type rule | t1, t2, … via using/typedef | fn tN() -> RustType with no arguments |
| Parameters | free-form (o, it, key, dst, n, …) | must be a0, a1, … consecutive from 0 |
| Generics | T1, T2, … (type and non-type params) | T1, T2, … consecutive from 1 |
| Variadic pack | typename... Args | trailing va: &[VaArg] |
| Constructing pack | Init<T, Args> &&...args | trailing init: RustType |
| Locals in Rust bodies | double-underscore prefix: __v, __fd, __e, … |
Notes:
- Rule numbering is per module, and gaps are currently allowed (e.g.
rules/maphas nof4), though this might change in the future. Names must be unique acrosssrc.candsrc.cppcombined. - On the C++ side parameter names are free, but the order defines the
placeholder indices: the first parameter is
a0on the Rust side, the second isa1, and so on. The receiver of a method rule is always the first parameter, hencea0. - Generic parameters are matched positionally between the two sides, so
T1in the Rust target means “whatever bound toT1in the C++ pattern”. - Locals introduced inside Rust rule bodies use a
__prefix. This is not checked by the build, but it is needed: rule bodies are spliced inline into the generated code, so an unprefixed local could collide with a variable name from the translated program.
Function qualifiers
- In
tgt_unsafe.rs, expression rules areunsafe fn; type rules (tN) are plainfn. - In
tgt_refcount.rs, all rules are safefn. The refcount model produces fully safe Rust, so a refcount rule body must not needunsafe.
The build does not check the qualifiers themselves; only rustc’s usual rules
apply when the rules crate compiles. In particular, nothing stops an unsafe
block inside a refcount rule body from being spliced into the output, so keeping
refcount rules safe is what upholds the model’s safety guarantee.
C++ pattern shape
- An
fNbody must be exactly onereturnstatement. The preprocessor rejects anything else. returnstatements are not allowed inside Rust rule bodies; write the result as a tail expression instead.- Exercise exactly one construct per rule. If an API has several overloads,
write one rule per overload (including separate rules for
const T &versusT &¶meters).
Argument accesses
Every use of an aN parameter in a rule body is classified as a read, write, or
move by the rule preprocessor. Passing
an argument by value counts as a read, not a move; the only way to record a move
is std::mem::take(&mut aN).
Type checking
All tgt_*.rs files are compiled as part of the rules crate, so a rule body
that does not type-check against libcc2rs, libc, nix, etc. breaks the
build. If a rule needs a new crate dependency, add it to rules/Cargo.toml and
to the hardcoded crate list in rule-preprocessor/src/semantic.rs (see
The Rule Preprocessors).
The Rule Preprocessors
Two build-time tools compile rule modules into the Rules IR that
cpp2rust loads at runtime. Both write into <build>/rules/<module>/:
cpp-rule-preprocessorcompilessrc.cpp/src.cintoir_src.json.rule-preprocessorcompilestgt_unsafe.rs/tgt_refcount.rsintoir_unsafe.json/ir_refcount.json.
The C++ side is keyed by resolved callee signatures; the Rust side by rule
names. The two are joined by rule name when cpp2rust loads them.
cpp-rule-preprocessor
A clang LibTooling executable (cpp2rust/cpp_rule_preprocessor.cpp) that runs
once per rule directory:
cpp-rule-preprocessor --dir rules/string --out <build>/rules/string/ir_src.json
Extra compiler flags can be passed with repeated --cxxflags options, though
CMake, which invokes the tool for every rule module via the
preprocess-cpp-rules target, passes none. Note that the parent directory of
--out must already exist; CMake creates it before each invocation, so a manual
run must do the same.
Rule sources are always compiled with the fixed flag set from
cpp2rust/compat/platform_flags.h, the same one used to parse input programs
(see Compat Shims). There is no compilation database and no
-std= flag: the language is chosen by clang from the file extension, and
src.c is processed before src.cpp.
For each rule it:
- Validates that every
fNbody is exactly onereturnstatement. - Resolves the callee of the returned expression. For non-template rules this is just the called declaration. For template rules the callee is unresolved, so the tool instantiates the rule’s template parameters with synthesized dummy types and runs overload resolution to find the function the rule refers to.
- Prints the resolved declaration as a canonical signature string:
<return type> <qualified::name>(<param types>[, ...])[ const][ volatile][ &|&&], where, ...appears for C-variadic functions and the trailing qualifiers only for methods. FortNaliases it prints the underlying type.
The output is a flat JSON object mapping rule names to these signature strings.
The printer preserves typedef sugar instead of desugaring it: size_t prints as
size_t, not unsigned long, which is what lets it map to usize while plain
unsigned long maps to u64 (for tN aliases this preservation is explicit;
inside function signatures the spelling survives through the printing policy).
Integer literals expanded from a macro are recorded as the macro name, which
is how constant rules
like the O_CREAT one match by name.
rule-preprocessor
A Rust binary crate built with the nightly toolchain because it links the
compiler’s own libraries (rustc_driver, rustc_middle, …). It processes the
whole rules tree in one invocation:
CARGO_TARGET_DIR=<target> cargo +nightly build --release \
--message-format=json-render-diagnostics \
--manifest-path rule-preprocessor/Cargo.toml > <target>/artifacts.json
RULE_PREPROCESSOR_ARTIFACTS=<target>/artifacts.json \
CARGO_TARGET_DIR=<target> cargo +nightly run --release \
--manifest-path rule-preprocessor/Cargo.toml -- <build>/rules [rules-dir]
The environment is load-bearing:
-
RULE_PREPROCESSOR_ARTIFACTSmust be set (the tool aborts otherwise): it points to cargo’s JSON build output, from which the rlibs of the rule dependencies (libcc2rs,libc,nix, …) are taken. The target dir is not scanned directly because it may hold stale copies of the same crates with different hashes (e.g., fromcargo clippyor an older toolchain), and mixing them makes type checking fail. The crate list is hardcoded, so a new dependency inrules/Cargo.tomlalso needs an entry inrule-preprocessor/src/semantic.rs. -
The sysroot comes from running
rustc --print=sysroot, so therustconPATHmust be the same nightly the preprocessor was built with (running throughcargo +nightly runguarantees this). -
rules-diris optional and defaults to the relative path../rules, resolved against the current working directory of the process.
CMake drives all of this via the preprocess-rust-rules target: it first builds
the rules crate with the stable toolchain (which also regenerates
rules/src/modules.rs), then builds the preprocessor in
<build>/rule-preprocessor-target, saving cargo’s output to artifacts.json
there, and runs it. That initial cargo build of the rules crate is what
actually gates the build on rule bodies type-checking (see below). The
preprocessor works in two phases.
Phase 1, syntactic. Each tgt_*.rs file is parsed with rust-analyzer’s
parser, and functions whose #[cfg] does not match the host are dropped. Every
function body is then turned into a list of fragments, whose kinds are
described in The Rules IR.
The fragmentation is mainly concerned with how the rule’s arguments are used:
references to parameters and generics become placeholder and generic fragments,
while source text that does not involve an argument is kept as-is.
Each placeholder is tagged with an access: read, write, or move. Some uses
give the access away syntactically (&mut a0 is a write); those that do not,
typically method-call receivers and arguments, are left as unknown for
phase 2. This phase also applies the two
preprocessor-side rewrites that
support rule rewriting.
Phase 2, semantic. The preprocessor compiles the rules crate in-process
with rustc and walks the typed HIR. This gives it the real signature of every
callee, which resolves the unknown accesses: passing to a &mut/*mut
parameter is a write, to a &/*const parameter a read, and to
std::mem::take a move. For type rules it also records which of the nine
derivable standard traits (Copy, Clone, Debug, Default, PartialEq,
Eq, PartialOrd, Ord, Hash) the mapped type implements. A placeholder
still unknown after this phase fails the build.
The preprocessor assumes the rules crate is buildable, which the earlier
cargo build of the crate ensures; errors from the in-process compilation are
therefore only reported as a warning.
The result is one ir_<model>.json per input file, keyed by rule name. The
output file name is derived from the input file name (tgt_unsafe.rs becomes
ir_unsafe.json) and the module directory is the direct parent of the
tgt_*.rs file.
The Rules IR
Each rule module compiles to up to three JSON files in
<build>/rules/<module>/:
ir_src.json: the C++ side, fromcpp-rule-preprocessor.ir_unsafe.json: the Rust side for the unsafe model, fromrule-preprocessor.ir_refcount.json: the Rust side for the refcount model, also fromrule-preprocessor(only if the module has atgt_refcount.rs).
All three are objects keyed by rule name (f1, t1, …), and the loader joins
them by name.
Source IR (ir_src.json)
A flat map from rule name to the canonical signature of the C++ construct the
rule matches. For rules/vector:
{
"t1": "std::vector<T1>",
"f3": "_Bool std::vector<T1>::empty() const"
}
This signature string is the lookup key for the whole rule: the converter prints C++ constructs from the input AST with the same printer and compares the strings.
Parameter packs
A function parameter pack prints as &&..., so the key of a rule for
emplace_back(Args &&...args) matches calls with any number of arguments. A
rule that takes an
init value also
records which template argument of the callee init builds:
"f112": {
"key": "T1 & std::vector<T1>::emplace_back(&&...)",
"init_type": { "depth": 0, "index": 0 }
}
Target IR (ir_unsafe.json / ir_refcount.json)
An expression rule serializes as an ExprRule object: the rule’s signature plus
its body as a list of fragments. For
unsafe fn f6<T1>(a0: &mut Vec<T1>) -> *mut T1 { a0.as_mut_ptr() }:
"f6": {
"body": [
{ "method_call": {
"receiver": [ { "placeholder": { "arg": 0, "access": "read" } } ],
"body": [ { "text": ".as_mut_ptr()" } ] } }
],
"generics": { "T1": [] },
"params": { "a0": { "type": "&mut Vec<T1>" } },
"return_type": { "type": "*mut T1", "is_unsafe_pointer": true }
}
The fragment kinds are:
text: literal Rust source, emitted verbatim.placeholder: a use of one of the rule’saNparameters in the body (not an argument of whatever the body calls); the converter substitutes the translated call-site argument here. Its fields:arg: the parameter index N.access: how the body uses the argument:read,write, ormove.is_index_base: the placeholder is the base of an index expression.
generic: aTNslot, replaced with the instantiated Rust type; serialized as the 1-based index N.method_call: a method call split intoreceiverandbodyfragment lists, so the code generator can rewrite the pair (see Rule Rewriting).va_args: the expansion point for a variadic tail.init: the value built from the call’s trailing arguments.
Every type in the Rules IR (in params, return_type, and type rules) is a
TypeInfo object, the type text plus a set of flags:
is_refcount_pointer: the type is aPtr<...>.is_unsafe_pointer: the type is a raw*mut/*constpointer.derives(type rules only): the standard traits the mapped type implements (Copy,Clone,Default, …).
The two pointer flags are mutually exclusive; the loader rejects a type with both set.
An ExprRule carries two flags of its own:
multi_statement: the body has more than one statement, or a statement followed by a tail expression, and must be wrapped in a block to stay a single expression.is_extern: the rule is an extern passthrough declaration and has no body.
Fields that are false, empty, or unset are omitted from the Rules IR. A va or
init parameter is never listed in params, and a () return type is omitted.
A type rule serializes as a TypeRule object: its TypeInfo plus the init
initializer expression, merged into one object:
"t1": { "type": "Vec<T1>", "init": "Default::default()" }
There is no explicit tag distinguishing the two rule kinds: an entry with body
is an expression rule, one with type and init a type rule.
In-memory form
cpp2rust mirrors the Rules IR in C++ structs of the same names, defined in
cpp2rust/converter/translation_rule.h.
TranslationRule::Load reads one module directory (the ir_*.json files
described above) and produces two maps keyed by rule name, one holding
ExprRules and one holding TypeRules:
- An
ExprRuleholds the body fragments, the parameter and returnTypeInfos, and the two rule-level flags (multi_statementandis_extern). The name-keyed Rules IR maps become positional vectors: parameteraNis entry N ofparams, genericTNis entry N-1 ofgenerics(each entry being the bound list). Rules support at most 9 generic parameters (kMaxGenerics). - A
TypeRuleholds the mapped type’sTypeInfoand theinitializerexpression. The same struct also represents the built-in type mappings (scalars, pointers, …) that the loader registers directly in code, without any Rules IR behind them: for example,intmaps toi32, andint *to*mut i32in the unsafe model orPtr<i32>in the refcount model.
Both also carry src, the canonical C++ signature attached from
ir_src.json; it is the key the rule is matched by.
How the loader finds the Rules IR directory, overlays the refcount model on the unsafe one, and indexes the loaded rules for matching is covered in Loading and Matching.
Loading and Matching
Finding the rules directory
cpp2rust takes the Rules IR directory via --rules <dir>. If the flag is
omitted it tries ./rules and then <executable dir>/../rules, accepting the
first candidate that contains (recursively) a subdirectory with ir_src.json
plus ir_unsafe.json or ir_refcount.json. Since the build writes the Rules IR
to <build>/rules and the binary lands in <build>/bin, the default resolution
picks up the generated Rules IR without any flags.
Loading
Rules are loaded once per process by Mapper::LoadTranslationRules:
- Built-in type mappings are registered first. Every scalar is mapped with its
width taken from the host (
intmaps toi32,unsigned longtou64), together with itsconstform and its pointer forms:*mut/*constin the unsafe model,Ptr<T>in the refcount model, where constness is dropped.charmaps tolibc::c_charin the unsafe model and tou8in refcount;size_t/ssize_tmap tousize/isize;void *maps to*mut ::libc::c_void, or in refcount toAnyPtr. - Every subdirectory of the rules directory is loaded with
TranslationRule::Load, which readsir_unsafe.json, overlaysir_refcount.jsonwhen translating with the refcount model, and then attaches the C++ signature fromir_src.jsonto each rule by name.
Loading is strict: an ir_src.json entry with no matching target rule is a
fatal error (this is what catches mismatched #if/#[cfg] gating), every
generic declared by a rule must appear in its C++ signature, and two type rules
for the same C++ type are rejected.
Matching
Loaded rules are indexed in two multimaps, one for expression rules and one for type rules. The multimap key is only a coarse bucket for collecting candidate rules; whether a candidate actually matches is decided by the matching engine. The bucket key is derived from the C++ signature:
- For expressions, the qualified function name with the return type, the
parameter list, and all template arguments stripped, so the rule for
_Bool std::vector<T1>::empty() constlands in thestd::vector::emptybucket. - For types, the text up to the first
<, so allstd::vector<...>rules land in thestd::vectorbucket.
During translation, the converter prints the construct it encounters with the
same canonical printer used by cpp-rule-preprocessor, which is what makes the
two sides comparable:
- Functions and methods print as
<return type> <qualified::name>(<param types>[, ...])[ const][ volatile][ &|&&]. - Enum constants and global variables print as their qualified name.
- Integer literals expanded from a macro print as the macro name.
All rules in the matching bucket are then unified against this string by
the matching engine, which binds T1…T9 to the concrete
types at the use site and picks the most specific rule when several match. Type
lookups first try the sugar-preserving spelling (so a rule can match size_t as
written) and retry with the desugared type on failure.
Running cpp2rust with --verbose logs every lookup and the rule it matched,
which is the quickest way to see why a rule does or does not fire.
Application
When a rule matches, the converter walks its body fragments and emits:
textfragments verbatim,placeholderfragments as the translated call-site argument. How the argument is emitted depends on the placeholder’s access and on whether the argument and the declared parameter type are pointers:- Read access emits the argument as a plain value, with an implicit numeric cast when the parameter type asks for one.
- Write access emits the argument as an lvalue.
- Move access wraps the argument in
std::mem::take(&mut ...); temporaries are moved as-is. - If the rule declares a pointer parameter but the argument is not a pointer,
the converter takes a fresh pointer to it,
materializing a temporary when the argument has
no address of its own. For example, the refcount
std::maxrule declaresPtr<T1>parameters, since the C++ side takesconst T1 &and the refcount model translates references asPtr, sostd::max(x1, x2)on plain locals substitutesx1.as_pointer()andx2.as_pointer()for the placeholders, whilestd::max(30, 40)first materializes__tmp_0and__tmp_1values for the literals and points into those. - If the receiver argument is a pointer but the rule expects a value, the converter dereferences it; this only happens for receivers, not for ordinary arguments.
genericfragments as the Rust mapping of the bound C++ type,va_argsfragments as the converted variadic tail,initfragments as the value built from the call’s trailing arguments,method_callfragments as receiver followed by body, possibly rewritten (see Rule Rewriting).
Multi-statement bodies are wrapped in { } so they remain a single expression.
Rules for user-defined C++ types are injected through the same mechanism at
translation time.
The Matching Engine
Loading and Matching collects the candidate rules for a
construct from a bucket; each candidate’s source signature is then unified
against the printed construct. The signature is treated as a template whose
T1…T9 slots capture concrete types: std::vector<T1>::vector() unifies
with std::vector<int>::vector() by binding T1 = int.
Unification works on the two strings:
- Whitespace differences are ignored.
- A
TNslot captures up to the next literal text of the pattern, found at the same<>/()/[]nesting depth. This is howT1captures all ofstd::map<int, int>instd::vector<std::map<int, int>>without stopping at the inner comma. - A
TNthat appears again must match its first capture exactly. - The whole printed string must be consumed; trailing text fails the match.
- Slots may stay unbound (a pattern can use
T2withoutT1).
A rule matches if unification succeeds. When several rules in the bucket match, the one with the longest source signature wins, so more specific rules take precedence; between equally long signatures the choice is unspecified.
Bucket keys
The bucket keys described in Loading and Matching have
two special cases: array types bucket by the text after the first [ rather
than the text before a <, and operator() rules are cut at the operator’s own
parentheses, so their key ends at ...::operator.
Instantiating the target
Captures are C++ spellings. Before being substituted into the rule’s Rust
fragments, each capture is itself mapped through the type rules, recursively, so
T1 = std::vector<int> substitutes as Vec<i32>. A captured type with no type
rule of its own is an error.
Rule Rewriting
A rule body is written against idiomatic Rust types: a rule that mutates a
vector declares its parameter as &mut Vec<T1>. But in the refcount model the
call-site argument is usually a Ptr<Vec<T1>>, and a Ptr
cannot produce a long-lived &mut. Instead of forcing
every rule to handle pointers, the code generator rewrites the rule body at
application time.
The with_mut rewrite
libcc2rs provides
#![allow(unused)]
fn main() {
impl<T> Ptr<T> {
pub fn with_mut<R>(&self, f: impl FnOnce(&mut T) -> R) -> R { ... }
}
}
which checks the pointer, borrows the pointee mutably, and runs the closure on
it (with an immutable sibling Ptr::with). The refcount converter uses it to
bridge the gap; the unsafe converter never rewrites and simply emits receiver
followed by body. The rewrite fires when all three hold:
- The rule body fragment is a method call whose receiver contains a placeholder (the preprocessor splits every method call into receiver and body fragments precisely to enable this). If the receiver contains several placeholders, the first one is used.
- The receiver placeholder’s access is write or move, i.e. the method takes
&mut selfor the rule mutates the parameter. Read access does not need the rewrite, since a read can go through aStrongPtrobtained withPtr::upgrade, or through aread()copy. - The call-site argument is a pointer, or an expression of reference type (which includes an operator call returning a reference).
The rule’s method call a0.method(...) is then emitted as
#![allow(unused)]
fn main() {
ptr.with_mut(|__v: <rule param type>| __v.method(...))
}
For example, the push_back rule is written as an ordinary &mut method call:
#![allow(unused)]
fn main() {
fn f21<T1: Clone>(a0: &mut Vec<T1>, a1: T1) { ... a0.push(...) }
}
Given the C++ input v.push_back(20); where v is reached through a
Ptr<Vec<i32>>, the generated code is:
#![allow(unused)]
fn main() {
v.with_mut(|__v: &mut Vec<i32>| __v.push(20));
}
When the receiver is a plain local value rather than a pointer, condition 3
fails and no closure is emitted; the same rule produces a direct call like
(*v2.borrow_mut()).push(0);.
The rewrite applies to pointer dereferences (p->push_back(20)) and to
reference usages (r.push_back(20) with std::vector<int> &r = *p); both are
translated as a Ptr, and that Ptr is what
with_mut is called on.
When the pointee is itself a boxed value (Value<T>, i.e. Rc<RefCell<T>>),
the closure takes &mut Value<T> and an extra borrow is inserted. This is the
case for nested containers: the refcount model translates
std::vector<std::vector<int>> as Vec<Value<Vec<i32>>> so that each element
has interior mutability of its own, and a Ptr to an inner vector therefore
points at a Value<Vec<i32>>, not a Vec<i32>:
#![allow(unused)]
fn main() {
ptr.with_mut(|__v: &mut Value<Vec<i32>>| (*__v.borrow_mut()).push(20))
}
The closure type is built from the C++ argument’s type, not the rule’s declared parameter type.
The read-access counterpart
For read access the converter does not emit a closure. A pointer receiver whose
rule parameter is a value or & type is simply dereferenced (p.read() or
(*p.upgrade().deref())); conversely, if the rule declares a Ptr parameter
but the argument is not a pointer, the converter inserts an as_pointer() cast
or materializes a temporary.
Preprocessor-side rewrites
Two rewrites in rule-preprocessor exist to make the with_mut rewrite
possible. Both apply only to &mut parameters:
- A
*deref in front of the parameter is dropped from the body, since the substituted argument is already an lvalue or pointer expression. std::mem::take(&mut aN)collapses to a bare placeholder, so the converter can re-express the move against the actual argument (for a pointer that becomesstd::mem::take(&mut <lvalue>)on the borrowed pointee). The spelling must be exactly this fully qualified form:mem::takeor an importedtakeis not rewritten. The collapsed placeholder’s access is leftunknownin phase 1; phase 2 resolves thestd::mem::takecall to a move.
Overview
Proving ownership in the presence of C++’s unrestricted aliasing is undecidable
in general, so cpp2rust does not try to satisfy Rust’s borrow checker
statically. Its default output, the refcount model, moves Rust’s ownership and
mutability checks to run time: reference counting replaces static ownership, and
dynamic borrow checks replace static mutability checks. This trades some speed
for safety, and it lets every program be translated.
libcc2rs is the runtime library where those checks live: a small crate of
auxiliary types and functions, such as Value<T>, Ptr<T>, and AnyPtr, that
translated programs link against. Keeping this machinery in one library keeps
the generated refcount code free of unsafe.
Every Rust file cpp2rust emits imports the whole crate:
#![allow(unused)]
fn main() {
extern crate libcc2rs;
use libcc2rs::*;
}
Module map
The modules fall into three groups.
The refcounted pointer model, the core of the refcount output:
rc:Value<T>andPtr<T>, the refcounted stand-ins for C values and pointers.cstr: string literals and thestring.hmemory functions overPtr<u8>byte strings.void:AnyPtr, the type-erased pointer forvoid *.ptr_dyn:PtrDyn<dyn T>, pointers to virtual classes.reinterpret: theByteReprtrait and allocation views that let a refcounted allocation be reinterpreted at the byte level, as C pointer casts do.alloc:malloc,free,realloc, andcallocover refcounted byte arrays.
Language-feature emulation, used by both models:
incanddec: traits implementing the four++/--operator forms.iterators: iteration for C++ containers that need stable iterators, with an implementation for both refcount and unsafe, and over C strings up to the null terminator.fn_ptr:FnPtr, function pointers with C-style address identity.va_args:VaArgandVaList, the representation of variadic calls.- The
goto,goto_block, andswitchproc macros, re-exported fromlibcc2rs-macros, which rewrite unstructured control flow into state machines.
The OS and libc surface:
io:CFilestreams, the standard streams, and read/write helpers.format:printf-style format string evaluation.fd: a registry tying integer file descriptors to their owning objects.- libc shims: safe wrappers over libc APIs, one
shim.rsper rule directory (files, directories, sockets, name resolution, polling, terminal control, time, and so on), compiled into the crate at build time. compat: platform-specific definitions, such as the location oferrnoandmalloc_usable_size.
Dependencies
The crate has five dependencies:
libcc2rs-macrosprovides the control-flow andderive(ByteRepr)proc macros.libcandnixprovide the raw and safe OS interfaces the shims wrap.jiffbacks the time shims.sprintfbacksprintf-style formatting.
Reference Counting
The refcount model produces safe Rust, and rc.rs is where that safety comes
from. It defines the two types every translated program is built on: Value<T>,
the translation of a C++ variable, and Ptr<T>, the translation of a C++
pointer.
Values and pointers
Rust requires every value to have a single owner, known at compile time, and
references to follow the borrow rules. C++ promises neither: a variable can be
aliased by any number of pointers, and any of them may write. Proving ownership
in the presence of such unrestricted aliasing is undecidable in general, so the
refcount model does not try. Instead it moves Rust’s ownership and mutability
checks from compile time to run time, trading some speed for safety: Rc counts
references and checks lifetimes dynamically, and RefCell checks at each access
that readers and writers do not overlap.
A C++ variable is therefore translated as a Value<T>, an alias for
Rc<RefCell<T>>. Taking the address of a variable becomes a call to
as_pointer, which produces a Ptr<T>:
int b = 2;
int *b_ptr = &b;
*b_ptr = 3;
#![allow(unused)]
fn main() {
let b: Value<i32> = Rc::new(RefCell::new(2));
let b_ptr: Value<Ptr<i32>> = Rc::new(RefCell::new(b.as_pointer()));
(*b_ptr.borrow()).write(3);
}
Weak references
A C++ pointer does not own what it points to, and Ptr<T> keeps that property:
it holds a Weak reference to the allocation plus an element offset. Ownership
stays with the variable binding for stack values and with the allocation itself
for the heap. When the owner goes away, every pointer into it dangles, and the
next access panics instead of reading freed memory.
The choice of weak over strong references is about destructors. C++ RAII code relies on destructors running at precise points, such as a mutex being released at the end of a scope; a strong reference held by a stray pointer could keep the object alive past that point and run its destructor late. With weak references, objects die exactly where C++ says they do, and a pointer that outlives its object dangles.
This is the central property of the model: memory bugs of the original program,
such as use after free, double free, and null or out-of-bounds dereference,
become panics in the translated one. Their messages carry the ub: prefix.
Pointer kinds
A Ptr<T> knows what it points into:
Null: the null pointer, and the default value.StackSingleandHeapSingle: a single value.StackArrayandHeapArray: a fixed-size array.StackVecandHeapVec: a growable buffer;std::vectorcontents and string literals live in one.Field: a field of a struct (see Pointers to fields).Reinterpreted: a byte-level view produced by a cast (see Type Reinterpretation).
An array carries one reference counter for the whole allocation, not one per element: the pointer pairs a weak reference to the whole array with the offset of the element it points to, which keeps the memory and performance overhead of arrays low.
Two pointers compare equal when they point into the same allocation at the same byte offset, and ordering compares allocation addresses, as C++ pointer comparison does.
Pointers to fields
Fields are stored inline in their struct, so a struct and all of its fields
share one allocation and one RefCell. A pointer to a field is to a struct what
a pointer to an element is to an array: it records a weak reference to the
allocation that holds the struct, its root, together with the byte offset of the
field in it. field_ptr! creates one:
struct point { int x; int y; };
struct point p;
int *y = &p.y;
#![allow(unused)]
fn main() {
let p: Value<point> = Rc::new(RefCell::new(<point>::default()));
let y: Value<Ptr<i32>> = Rc::new(RefCell::new(field_ptr!(p, y)));
}
field_ptr!(p, y) works on a Value or a Ptr to a struct. The offsets are
those of the C layout of the struct, which the code generator gets from Clang
and writes on the fields as #[offset(N)] attributes (any constant expression
works, e.g., offset_of! in the libc shims);
#[derive(Record)] turns them into the implementation of the Record trait:
#![allow(unused)]
fn main() {
#[derive(Clone, Record, Default)]
pub struct point {
#[offset(0)]
pub x: i32,
#[offset(4)]
pub y: i32,
}
}
Creating a field pointer allocates nothing. The byte offset of the field is kept in the pointer’s offset, which for a field pointer counts bytes, as for a reinterpreted pointer. Pointers to fields of fields, and to fields of array elements, add up the offsets, and keep the allocation of the outermost struct or array as their root.
An access through a field pointer borrows the root, and finds the field from its
offset and its type with Record::locate, which the derive generates as a
comparison of the offset against those of the fields, recursing into nested
structs. A field is identified by its type too, as a struct and its first field
share an offset. An offset that doesn’t find a field of the pointer’s type
panics with ub: invalid field pointer.
Arrays and vectors are not stored inline: an array field is a Value<Box<[T]>>
of its own, and a std::vector field a Value<Vec<T>>. Pointers to their
elements are ordinary array pointers, and pointers to the fields of those have
the array or vector as their root. A field pointer hence always points to a
single object, which is why it needs no element index.
array_field_ptr!(p, arr) is a pointer to element 0 of the array field arr,
and is how the elements of an array field are accessed. It is an ordinary array
pointer into the array’s Value, except when p is a reinterpreted pointer:
the array then lives in the bytes of the original allocation, so the result is a
reinterpreted pointer to those bytes, at the offset of the field. Writes to the
elements hence reach the original allocation, and the pointer can go past the
end of the struct, as C code does with a trailing char name[1] in an
over-allocated struct.
Reading a field through a pointer borrows the struct only for the duration of a
closure, e.g., p.with(|s| s.y), so that the borrow ends before the rest of the
statement runs. field!(p, y) is the place of the field, which is read and
written through p, projected to the field by a closure, without looking the
field up: field!(p, y).write(3). It holds a reference to p, and takes no
more room than a pointer. field_ptr! is only used where an actual pointer to
the field is needed.
A field of a reinterpreted struct is itself a reinterpreted pointer, to the bytes of the field. Conversely, reinterpreting a field pointer views the bytes of the field alone, as if it were its own allocation.
Since all fields share the borrow of their struct, an expression must not write to a field while another field of the same struct is borrowed. The code generator scopes the borrows of field reads so that they end before any write in the same statement.
The heap
new and new[] are translated as Ptr::alloc and Ptr::alloc_array, and
malloc, calloc, realloc, and strdup allocate through Ptr::alloc_array
as well; the alloc module defines them as named functions (malloc_refcount,
free_refcount, realloc_refcount, calloc_refcount, strdup_refcount, and
their _unsafe twins for the unsafe model). The allocation’s Rc is leaked
with Rc::into_raw so the object outlives the statement that created it, and
delete and delete_array recover the leaked reference and drop it:
int *d = new int(0);
*d = 5;
delete d;
#![allow(unused)]
fn main() {
let d: Value<Ptr<i32>> = Rc::new(RefCell::new(Ptr::alloc(0)));
(*d.borrow()).write(5);
(*d.borrow()).delete();
}
delete checks that the pointer still points at the start of a live heap
allocation: freeing twice, freeing through an offset pointer, or freeing a stack
or Vec pointer panics with ub:.
A heap allocation can also change hands instead of being freed. to_owned_opt
recovers the leaked reference the same way delete does, but returns it to the
caller as an owning Option<Value<T>> (or Option<Value<Box<[T]>>> for an
array), with None for the null pointer; from then on the allocation lives
exactly as long as that binding. It panics for stack, Vec, and reinterpreted
pointers. This is how std::unique_ptr<T> is translated: the smart pointer is
an Option<Value<T>>, its constructor and reset adopt a raw pointer with
to_owned_opt, and as_pointer, which is also implemented for
Option<Value<T>> and yields null for None, stands in for get():
std::unique_ptr<int> u(new int(1));
int *raw = u.get();
#![allow(unused)]
fn main() {
let u: Value<Option<Value<i32>>> =
Rc::new(RefCell::new(Ptr::alloc(1).to_owned_opt()));
let raw: Value<Ptr<i32>> = Rc::new(RefCell::new((*u.borrow()).as_pointer()));
}
Dereferences
A dereference becomes a short-lived borrow. read and write copy a value out
of or into the allocation:
*d = 5;
int v = *d;
#![allow(unused)]
fn main() {
(*d.borrow()).write(5);
let v: Value<i32> = Rc::new(RefCell::new((*d.borrow()).read()));
}
A Ptr cannot simply return a &T or &mut T to its pointee: the reference
would keep the RefCell borrowed with nothing to bound its lifetime. with and
with_mut invert the control instead: the expression that needs the reference
moves into a closure, and the borrow lasts exactly as long as the closure runs.
They carry the operations that need a reference to the existing value, such as a
push_back on a vector reached through a pointer (write could only replace
the vector wholesale):
#![allow(unused)]
fn main() {
v.with_mut(|v| v.push(20));
}
Applied rule bodies are the main producer of these calls (see
Rule Rewriting). write itself is a thin wrapper: it
is defined as with_mut(|v| *v = value).
with_slice and with_slice_mut are the same idea over a range of elements:
they expose len bytes starting at the pointer as a Rust slice for the duration
of a closure, which is how a C buffer is passed to Rust and nix functions such
as read and write.
In every case the RefCell is borrowed only for the duration of the access,
which is what lets freely aliasing C++ pointers coexist with the borrow checker:
no borrow outlives the expression that created it. When an expression needs an
actual Rust reference, the pointer is upgraded to a
StrongPtr, which holds the allocation alive and hands out
a Ref.
These borrows are the model’s mutability checks, moved from compile time to run
time. Rust’s rule still holds, any number of readers or one writer, but it is
enforced when the access happens: an expression that writes a variable while
also reading it through an alias, such as *x.borrow_mut() = *x.borrow() + 1,
traps. The code generator is responsible for not emitting such expressions: it
stores intermediate results in temporaries, so the reading borrow ends before
the writing borrow starts.
Strong pointers
upgrade turns a Ptr<T> into a StrongPtr<T>, the same pointer holding a
strong Rc to its allocation instead of a weak one:
#![allow(unused)]
fn main() {
pub enum StrongPtr<T> {
StackSingle(Rc<RefCell<T>>),
Vec { rc: Rc<RefCell<Vec<T>>>, offset: usize },
StackArray { rc: Rc<RefCell<Box<[T]>>>, offset: usize },
Field { root: Rc<dyn Root>, offset: usize },
Reinterpreted { alloc: OriginalAlloc, byte_offset: usize, cell: RefCell<Option<T>> },
}
}
deref returns a Ref<'_, T> to the pointee. The Ref borrows the
StrongPtr, so the borrow of the RefCell lasts as long as the strong pointer
does: in p.upgrade().deref().field, the temporary StrongPtr lives until the
end of the enclosing statement, and so does the borrow. The code generator
prefers the with and with_mut closures, and upgrades only where the result
of an access must borrow the pointee beyond a closure. There is no deref_mut;
writes go through write and with_mut, including those to a field of a struct
reached through a pointer: field!(p, x).write(v).
For the Reinterpreted variant there is no value to reference, only bytes in
another allocation. deref reads those bytes into a local cell and hands out a
Ref to that copy, refreshing it on every call. with_mut writes the bytes
through to the original allocation before it returns, so that a write is visible
through every other pointer at once.
Warning
StrongPtris set to be removed. Holding a strong reference, even briefly, undermines the model in two ways:
- Nothing prevents a
StrongPtrfrom outliving its statement. One that is stored, returned, or bound to a local keeps the object alive past the point where C++ destroys it, so its destructor runs late and dangling accesses go unnoticed; a heap object held this way makes the laterdeletepanic. The generator only ever emits it as a temporary, but hand-edited code or a rule can break that.- Even as a temporary, it lives for the whole statement. In
(*p.upgrade().deref()).method()the strong reference is alive during the call, so a method that runsdelete this, or otherwise deletes the object it was called on, hitsdelete’s reference-count check and panics with a spuriousub: invalid delete.
Where the code generator needs a different view of the same allocation, it does
not upgrade at all. decay turns a pointer to a whole Vec<T> or Box<[T]>
into a pointer to its first element by re-tagging the existing weak reference,
and Ptr::to_dyn (see Virtual Classes) does the same for the
upcast to a trait object.
Arithmetic
The offset lives in the pointer, so arithmetic never touches the allocation.
p + n, p - n, and the ++/-- forms move the offset, including past the
end of the allocation, exactly as C++ allows; bounds are checked only when the
pointer is dereferenced. Subtracting two pointers yields their element distance
and requires both to point into the same allocation.
The pointer also knows the extent of its allocation: len is the number of
elements in it, whatever the pointer’s offset, and get_offset is the pointer’s
element index within it. The container rules build end() and back() pointers
from these (to_end, to_last) and turn a [first, last) range into a count
or an absolute index the same way.
Integer casts
Casts between pointers and integers are translated as to_int and from_int:
uintptr_t n = (uintptr_t)p;
int *q = (int *)n;
#![allow(unused)]
fn main() {
let n: Value<usize> = Rc::new(RefCell::new((*p.borrow()).to_int()));
let q: Value<Ptr<i32>> =
Rc::new(RefCell::new(<Ptr<i32>>::from_int(*n.borrow())));
}
Both currently panic when executed. Giving them well-defined semantics is work in progress (#225).
C Strings
C and C++ strings are byte strings: programs manipulate individual bytes and the
contents need not be valid UTF-8, so strings are translated as u8 buffers
rather than Rust String values. A string literal becomes a per-thread interned
buffer with a trailing zero byte, handed out as a Ptr<u8> by
Ptr::from_string_literal. Ptr<u8> also carries the memory functions C
strings rely on: memcpy (with memmove semantics for overlapping buffers
instead of undefined behavior), memset, memcmp, and to_rust_string for
crossing into Rust APIs.
CStringIterator, returned by to_c_string_iterator, walks the bytes of a
Ptr<u8> up to the null terminator; the string.h rules are built on it, and
Display for Ptr<u8> prints it, so a C string can be formatted directly.
Copying a string does not go through the iterator, which would grow its
destination as it goes. with_c_str scans for the null terminator in the
backing storage and lends the bytes to a closure, so measuring or borrowing a
string allocates nothing (c_str_len, and count on a CStringIterator, are
built on it). to_c_bytes and to_rust_string, and through them strdup and
the std::string constructors, copy the string with a single allocation.
to_c_bytes leaves room for one more byte, which callers usually spend on the
terminator or a newline.
Iterating a Ptr<T> reports how many elements are left, so collecting a Ptr
into a Vec also allocates once.
void Pointers
void * is translated as AnyPtr, a type-erased Ptr. to_any erases the
element type and reinterpret_cast recovers it:
char data[] = "hi";
void *vp = data;
char *cp = vp;
#![allow(unused)]
fn main() {
let data: Value<Box<[u8]>> = Rc::new(RefCell::new(Box::from(*b"hi\0")));
let vp: Value<AnyPtr> =
Rc::new(RefCell::new((data.as_pointer() as Ptr<u8>).to_any()));
let cp: Value<Ptr<u8>> =
Rc::new(RefCell::new((*vp.borrow()).reinterpret_cast::<u8>()));
}
reinterpret_cast returns the original pointer when the requested type matches
the erased one, and a byte-level view otherwise, because C
code commonly casts A * to void * and reads it back as B *.
The malloc family allocates and frees through AnyPtr, so the returned
pointer is cast to the requested type and cast back to free it:
int *p = malloc(sizeof(int));
*p = 42;
free(p);
#![allow(unused)]
fn main() {
// malloc_refcount(n) is
// Ptr::alloc_array(vec![0u8; n].into_boxed_slice()).to_any()
let p: Value<Ptr<i32>> = Rc::new(RefCell::new(
malloc_refcount(::std::mem::size_of::<i32>()).reinterpret_cast::<i32>(),
));
(*p.borrow()).write(42);
free_refcount((*p.borrow()).to_any());
}
AnyPtr also carries memcpy, memset, and memcmp, forwarding to the
Ptr<u8> versions from C Strings over the byte view of its
pointee.
Warning
Two
AnyPtrvalues are equal only when they were erased from the same pointer type and compare equal as that type. Avoid *obtained from aPtr<i32>and one obtained from aPtr<u8>into the same allocation compare unequal, where C would consider them the same address. This is set to be fixed by comparing through the byte view instead.
Casts between AnyPtr and integers use the same to_int and from_int as
Ptr<T>.
Virtual Classes
A pointer to a virtual class cannot be a Ptr<T>. Ptr<T> requires T: Sized,
the implicit bound on every generic parameter, and a virtual class is translated
as a Rust trait, whose trait object dyn T is unsized and cannot satisfy that
bound. The runtime provides a dedicated PtrDyn<dyn T> type, declared with
T: ?Sized, for these pointers, kept separate so the generic Ptr pays no cost
for dynamic dispatch.
A PtrDyn is created at the point where C++ converts a derived pointer to a
base pointer. Ptr::to_dyn takes the weak reference out of the Ptr<Derived>
and applies Rust’s unsized coercion to it, turning a Weak<RefCell<Derived>>
into a Weak<RefCell<dyn Base>>, without ever upgrading it. The coercion itself
is written by the code generator as the closure |w| w, whose return type
selects the target trait object.
Note
This coercion is why the pointee cell is a plain
RefCellbehindWeakrather than a struct of the runtime’s own:Weakalready implements it, and a new type could only opt in through the nightly-onlyCoerceUnsizedtrait.
The conversion looks like this:
struct Base { virtual int f() const = 0; };
struct Derived : Base { int f() const override { return 1; } };
Derived d;
Base *b = &d;
int r = b->f();
#![allow(unused)]
fn main() {
let d: Value<Derived> = Rc::new(RefCell::new(<Derived>::default()));
let b: Value<PtrDyn<dyn Base>> = Rc::new(RefCell::new(
(d.as_pointer()).to_dyn::<dyn Base>(|w| w),
));
let r: Value<i32> = Rc::new(RefCell::new(
({ (*(*b.borrow()).upgrade().deref()).f() }),
));
}
A virtual call goes through upgrade, which returns a StrongPtrDyn<dyn T>
holding the strong reference; its deref and deref_mut borrow the object and
the call dispatches through the trait’s vtable.
Warning
StrongPtrDynis set to be removed for the same reasons asStrongPtr: it holds a strong reference that can outlive the object’s C++ lifetime, and even as a temporary it spans the whole virtual call, so a method that deletes its own object panics ondelete.
PtrDyn is far smaller than Ptr: it is either null or a weak reference to a
single object, on the stack or on the heap. Ptr::to_dyn keeps that
distinction, so a base pointer made from a newed object can be deleted and
one made from a local cannot. It has no arithmetic, no comparison, no array
kinds, and no byte view. Because to_dyn is only defined for single-value
pointers, a base pointer into an array of polymorphic objects (a
Derived arr[N] walked through a Base *) cannot be formed.
Type Reinterpretation
C code reads the same memory at different types: a long is inspected byte by
byte through a char *, a byte buffer from malloc is used as an array of
structs, a struct sockaddr_in is passed where a struct sockaddr is expected.
In the refcount model there are no raw bytes to point at: values are typed Rust
data behind refcounted cells. The reinterpret module supplies the byte view
these programs expect.
ByteRepr
ByteRepr gives a type its C byte representation:
#![allow(unused)]
fn main() {
pub trait ByteRepr: 'static {
fn byte_size() -> usize;
fn to_bytes(&self, buf: &mut [u8]);
fn from_bytes(buf: &[u8]) -> Self;
}
}
The code generator emits the ByteRepr implementation of a C struct next to it:
struct header {
int tag;
int size;
};
#![allow(unused)]
fn main() {
pub struct header {
pub tag: i32,
pub size: i32,
}
impl ByteRepr for header {
fn byte_size() -> usize {
8
}
fn to_bytes(&self, buf: &mut [u8]) {
self.tag.to_bytes(&mut buf[0..4]);
self.size.to_bytes(&mut buf[4..8]);
}
fn from_bytes(buf: &[u8]) -> Self {
Self {
tag: <i32>::from_bytes(&buf[0..4]),
size: <i32>::from_bytes(&buf[4..8]),
}
}
}
}
byte_size is sizeof(struct header). to_bytes writes each field at its C
offset into an 8-byte buffer, so the buffer holds the struct exactly as it would
sit in C memory. from_bytes reads such a buffer back into a fresh struct.
The primitive types serialize to their native-endian bytes, matching what C sees
on the host. The libc shims implement the
trait by hand with the byte layout of their C structs. Types with no meaningful
C layout, such as std::fs::File or Vec<T>, implement the trait with defaults
that panic, so reinterpreting one is caught at run time.
derive(ByteRepr)
libcc2rs-macros provides a #[derive(ByteRepr)] proc macro. It is implemented
only for unit structs, and expanding it on a struct with fields, an enum, or a
union is a compile-time error. The expansion sets byte_size to 1 and leaves
to_bytes and from_bytes on the trait’s panicking defaults:
#![allow(unused)]
fn main() {
#[derive(Default, Clone, Copy, ByteRepr)]
pub struct UnitStruct;
}
Views over the original allocation
reinterpret_cast copies nothing. It produces a Ptr in the Reinterpreted
kind: a handle to the original allocation plus a byte offset, stepping by the
target type’s size. A read serializes the overlapping elements of the original
into bytes and parses the target value out of them; a write is a
read-modify-write back into the original. The data always lives in the original
allocation, so writes through the original are visible through every view and
writes through a view are visible everywhere else:
#![allow(unused)]
fn main() {
let p: Ptr<u64> = Ptr::alloc(0x0807060504030201);
// A view over p's allocation at byte offset 0, stepping by 1 byte.
let bytes: Ptr<u8> = p.reinterpret_cast::<u8>();
// Read: p.to_bytes() gives the 8 bytes, u8::from_bytes parses byte 0.
assert_eq!(bytes.read(), 0x01);
// Write: p.to_bytes(), replace byte 7 with 0xAA, u64::from_bytes back into p.
bytes.offset(7).write(0xAA);
// The write went into the original allocation.
assert_eq!(p.read(), 0xAA07060504030201);
}
A reinterpreted pointer counts its offset in bytes, so its arithmetic matches the C cast exactly. Casting a view again does not stack views: the new pointer keeps the handle to the original allocation.
Deleting through a reinterpreted pointer frees the original allocation. That is
how free works on a buffer that has been cast around: the pointer is
reinterpreted to bytes and the original allocation is deleted.
Cost
A cast allocates exactly one small object, the view, which holds a weak
reference to the original allocation, the size of the target type, and a
reference to a stateless table of the byte-level operations for the original’s
storage type (a single value, a Vec, a boxed slice, or a field of a struct,
whose bytes are found through the struct’s allocation as for a
field pointer). Copying or offsetting the
resulting Ptr only bumps a reference count, so a loop over a malloced array
pays for the cast once, not per access.
Accessing memory through a view does not allocate: the bytes of the accessed
element are staged in a stack buffer (a heap buffer is used only for accesses
larger than 64 bytes). When the original allocation is a u8 buffer, as it is
for everything that comes from malloc, reads and writes copy the bytes
directly instead of serializing the elements one by one.
Known limitations
Reading a struct through a reinterpreted pointer builds a fresh struct with
from_bytes, which exists only for the duration of the with closure, or as
long as the StrongPtr that holds it. Writes to its fields go through
with_mut, which encodes the struct back into the original allocation before it
returns, so they are visible right away through any other pointer; a pointer to
one of its fields is a reinterpreted pointer into the original allocation.
However, union accessors return pointers to the union’s
storage; on a reinterpreted union that storage is the temporary, so the returned
pointer dangles. This is set to be fixed in the near future.
AnyPtr casts
AnyPtr::reinterpret_cast first tries to recover the pointer as it was erased:
casting a void * back to the type it came from returns the original Ptr<T>,
with no byte view involved. Only a cast to a different type goes through the
byte representation:
#![allow(unused)]
fn main() {
let p: Ptr<u64> = Ptr::alloc(0x0807060504030201);
let any: AnyPtr = p.to_any();
// Same type as erased: the original Ptr<u64> comes back.
let back: Ptr<u64> = any.reinterpret_cast::<u64>();
assert!(back == p);
// Different type: a byte view over p's allocation, as with
// Ptr::reinterpret_cast.
let bytes: Ptr<u8> = any.reinterpret_cast::<u8>();
assert_eq!(bytes.read(), 0x01);
}
Increment and Decrement
C’s ++ and -- are expressions: x++ yields the old value and ++x the new
one, and both can appear inside a larger expression. Rust only has the x += 1
statement, so the inc and dec modules define one trait per operator form:
#![allow(unused)]
fn main() {
pub trait PostfixInc { fn postfix_inc(&mut self) -> Self; }
pub trait PrefixInc { fn prefix_inc(&mut self) -> Self; }
pub trait PostfixDec { fn postfix_dec(&mut self) -> Self; }
pub trait PrefixDec { fn prefix_dec(&mut self) -> Self; }
}
Each method updates the value in place and returns what the C expression evaluates to: the postfix forms return a copy of the old value, the prefix forms the new one.
int x = 0;
while (x++ < 100 && x != 50) {
++x;
}
#![allow(unused)]
fn main() {
let x: Value<i32> = Rc::new(RefCell::new(0));
while (*x.borrow_mut()).postfix_inc() < 100 && *x.borrow() != 50 {
(*x.borrow_mut()).prefix_inc();
}
}
The traits are implemented for the integer types with wrapping arithmetic, so
overflow behaves as C’s unsigned wraparound and never panics, and for f32 and
f64. Ptr<T> implements them by moving its offset one
element, and the map iterators by stepping to
the neighbouring key. For each translated enum the code generator emits
impl_enum_inc_dec!, a macro exported by inc that implements the four traits
by converting through i32.
The unsafe model uses the same traits for integers and floats. For raw pointers
the same method names come from separate Unsafe* traits (UnsafePrefixInc and
so on), whose methods are unsafe fn and step the pointer with offset(1), so
the generated code reads the same in both models:
#![allow(unused)]
fn main() {
let mut q: *mut i32 = p;
q.prefix_inc();
q.postfix_dec();
}
Iterators
A C++ iterator is a pointer-like object: it is dereferenced, compared against
end(), and moved with ++ and --, and it stays usable across the statements
of a loop body. Rust iterators are consumed by a for loop and cannot be
compared or stepped backwards, so the runtime represents C++ iterators with
values of its own.
Random access iterators
For std::vector, std::string, and arrays the iterator is a
Ptr<T> into the container’s buffer: begin() is as_pointer(),
end() is to_end(), and comparison and arithmetic are the pointer’s own.
Ptr<T> also implements Iterator, yielding a pointer to each element, so a
range-based for becomes a Rust for over the pointer:
std::vector<int> v;
for (auto x : v)
printf("%d\n", x);
#![allow(unused)]
fn main() {
let v: Value<Vec<i32>> = Rc::new(RefCell::new(Vec::new()));
for x in v.as_pointer() as Ptr<i32> {
println!("{}", x.read());
}
}
Two variants serve special cases. StringIterator, returned by
to_string_iterator, stops before the trailing zero byte, so iterating a
std::string visits its characters only. PtrValueIter yields copies of the
elements instead of pointers to them; rule bodies use it to feed a range of C
memory to Rust iterator adaptors:
#![allow(unused)]
fn main() {
// std::accumulate(first, last, init)
let count = (last - first) as usize;
PtrValueIter::new(&first, count).fold(init, |acc, x| acc + x)
}
Stable iterators
std::map<K, V> is translated as a BTreeMap<K, Value<V>>, which has no
addressable elements to point into. The runtime defines MapIter for it: a pair
of a handle to the map and the current key, with None standing for end().
Because it stores a key rather than a position, it survives insertions and
removals elsewhere in the map, as C++ guarantees. begin, end, and find_key
construct one; inc and dec move to the neighbouring key; erase removes the
current entry and returns the iterator to the next; the ++/-- traits and
Iterator are implemented on top of these. Two iterators compare equal when
they hold the same key, so it != m.end() compares Some(key) against None:
std::map<int, double> m;
for (const auto &i : m)
sum += i.second;
#![allow(unused)]
fn main() {
let m: Value<BTreeMap<i32, Value<f64>>> =
Rc::new(RefCell::new(BTreeMap::new()));
for i in RefcountMapIter::begin(m.as_pointer()) {
(*sum.borrow_mut()) += (*i.second().borrow());
}
}
Warning
Equality only looks at the key: iterators into two different maps compare equal when they hold the same key. In C++ comparing them is undefined behaviour, so the translation should panic with
ub:instead; this will be fixed by comparing the map handles as well.
first() and second() come from the MapIterator trait and take the place of
it->first and it->second. MapIter is generic over how the map is reached,
which is what gives it an implementation for both models:
RefcountMapIter<K, V> holds a Ptr<BTreeMap<K, Value<V>>> and returns
Value<K> and Value<V>; UnsafeMapIterator<K, V> holds a
*const BTreeMap<K, Box<V>> and returns *const K and *mut V.
Function Pointers
A C function pointer can be null, compared for equality, cast to another
function pointer type and back, and stored in a void *. A Rust fn value can
be called and compared, but it is never null and its type is fixed, so the
refcount model translates function pointers as FnPtr<T>, where T is the Rust
fn type of the target:
#![allow(unused)]
fn main() {
pub struct FnPtr<T> { /* the function as first stored, and its current cast */ }
impl<T> FnPtr<T> {
pub fn null() -> Self;
pub fn new(f: T) -> Self;
pub fn is_null(&self) -> bool;
pub fn cast<U>(&self, adapter: Option<U>) -> FnPtr<U>;
pub fn to_any(&self) -> AnyPtr;
}
}
FnPtr dereferences to the function, so a call through it is (*fp)(args).
Calling a null pointer panics with ub:.
FnPtr stores the function inline, together with its address, which is how
pointers are compared; the FnAddr trait provides the address. Creating,
copying, and calling a function pointer does not allocate. Rust has no way to
write an impl for every fn arity at once, so FnAddr is implemented by a
macro for fn types of zero to sixteen parameters. A function with more
parameters cannot be wrapped in an FnPtr, and taking its address fails to
compile with a missing FnAddr bound.
typedef int (*int_fn)(int);
int double_it(int x) { return x * 2; }
int_fn fn = double_it;
int r = fn(5);
#![allow(unused)]
fn main() {
let fn_: Value<FnPtr<fn(i32) -> i32>> =
Rc::new(RefCell::new(FnPtr::<fn(i32) -> i32>::new(double_it_0)));
let r: Value<i32> = Rc::new(RefCell::new((*(*fn_.borrow()))(5)));
}
Casts
C code casts function pointers to a different type and calls through the new
type. When the two types are not compatible this is undefined behavior, but the
argument types involved usually have the same representation, so implementations
accept the call and programs rely on it. Below, add_offset takes an int *,
but is called through a pointer that takes a void *:
typedef int (*generic_int_fn)(void *, int);
int add_offset(int *base, int offset) { return *base + offset; }
generic_int_fn gfn = (generic_int_fn)add_offset;
int result = gfn(&val, 42);
In Rust fn(Ptr<i32>, i32) -> i32 and fn(AnyPtr, i32) -> i32 are unrelated
types, so the code generator emits an adapter: a function of the target type
that converts the arguments and calls the original. cast stores it, and calls
through the cast pointer go through the adapter:
#![allow(unused)]
fn main() {
let gfn: Value<FnPtr<fn(AnyPtr, i32) -> i32>> = Rc::new(RefCell::new(
FnPtr::<fn(Ptr<i32>, i32) -> i32>::new(add_offset_4)
.cast::<fn(AnyPtr, i32) -> i32>(Some(
(|a0: AnyPtr, a1: i32| -> i32 {
add_offset_4(a0.reinterpret_cast::<i32>(), a1)
}) as fn(AnyPtr, i32) -> i32,
)),
));
let result: Value<i32> = Rc::new(RefCell::new(
(*(*gfn.borrow()))(val.as_pointer().to_any(), 42),
));
}
The code generator can build an adapter when the arguments and return type of
the two function types have the same representation. Otherwise it passes None,
and calling through the cast pointer panics with ub:.
A cast to a different type is the only operation that allocates: the pointer then also keeps the function it was created with, type-erased, so that casting back to that type can restore it. Equality compares the address of the function the pointer was created with.
Casting a function pointer to void * is to_any, and AnyPtr::cast_fn::<T>
recovers it. reinterpret_cast on an AnyPtr holding a function currently
panics, as do integer casts on a Ptr; both are set to
be fixed in the near future.
Lambdas
A lambda is an FnPtr too, in both models (see
Lambdas). One without captures is built with
new from a closure, like a function. One with captures is built by the
lambda! and lambda_unsafe! macros, which declare a struct holding the
captures, with the body as its method, and pass both to from_lambda or
from_lambda_unsafe:
#![allow(unused)]
fn main() {
impl<A, R> FnPtr<fn(A) -> R> {
pub fn from_lambda<L>(lambda: L, call: fn(&L, A) -> R) -> Self;
pub fn from_lambda_unsafe<L>(lambda: L, call: fn(&mut L, A) -> R) -> Self;
}
}
The unsafe model translates function pointers as Option<unsafe fn> and uses
FnPtr only for lambdas. For this, FnPtrArg is also implemented for raw
pointers and for Option<unsafe fn>, and derived by the structs and unions of
the unsafe model.
Variadic Functions
Rust has no ... parameters and no va_list. A variadic C function is
translated as a function whose last parameter is a slice of VaArg, an enum
with one variant per kind of value C’s default argument promotions can produce:
#![allow(unused)]
fn main() {
pub enum VaArg {
Int(i32),
UInt(u32),
Long(i64),
ULong(u64),
Double(f64),
RawPtr(*mut c_void),
Ptr(AnyPtr),
}
}
At a call site every extra argument is converted with .into(), which performs
the promotions (char and short to int, float to double) and erases
pointers to AnyPtr in the refcount model or *mut c_void in the unsafe model.
Inside the function, va_list is a VaList, a cursor over the slice:
va_start becomes VaList::new(__args), va_arg(ap, T) becomes
ap.arg::<T>(), va_copy is a plain copy of the cursor, and va_end is a
no-op:
int sum(int count, ...) {
va_list ap;
va_start(ap, count);
int total = 0;
for (int i = 0; i < count; i++)
total += va_arg(ap, int);
va_end(ap);
return total;
}
sum(3, 10, 20, 30);
#![allow(unused)]
fn main() {
pub fn sum_0(count: i32, __args: &[VaArg]) -> i32 {
let ap: Value<VaList> = Rc::new(RefCell::new(VaList::default()));
(*ap.borrow_mut()) = VaList::new(__args);
let total: Value<i32> = Rc::new(RefCell::new(0));
// ...
(*total.borrow_mut()) += (*ap.borrow_mut()).arg::<i32>();
// ...
}
sum_0(3, &[10.into(), 20.into(), 30.into()]);
}
arg::<T>() goes through the VaArgGet trait, implemented for the integer and
floating types, raw pointers, Ptr<T>, AnyPtr, and FnPtr<T>. Integer
variants convert freely among the integer types, as va_arg does with types of
the same rank; asking for a pointer where an integer was passed, or the reverse,
panics, as does reading past the last argument.
Variadic libc functions such as printf and fcntl are handled by
variadic rules, whose bodies
receive the same &[VaArg] slice; format_c in the
format module consumes one to evaluate a format string.
Control Flow Macros
Rust has no goto, and a match arm never falls into the next one. The
libcc2rs-macros crate provides two procedural macros, re-exported by
libcc2rs, that express these C constructs as a state machine. Both models use
them.
goto_block
goto_block! takes a sequence of labeled blocks. Execution starts in the first
block and falls through from each block into the next; goto!('label) jumps to
the block with that label, forwards or backwards:
int retry(int n) {
int count = 0;
int acc = 0;
again:
count += 1;
acc += n;
if (count < 3)
goto again;
return acc;
}
#![allow(unused)]
fn main() {
pub fn retry_0(n: i32) -> i32 {
let n: Value<i32> = Rc::new(RefCell::new(n));
let count: Value<i32> = <Value<i32>>::default();
let acc: Value<i32> = <Value<i32>>::default();
goto_block!({
'entry: {
*count.borrow_mut() = 0;
*acc.borrow_mut() = 0;
}
'again: {
(*count.borrow_mut()) += 1;
(*acc.borrow_mut()) += (*n.borrow());
if *count.borrow() < 3 {
goto!('again);
}
return (*acc.borrow());
}
});
panic!("ub: non-void function does not return a value")
}
}
The code generator puts the statements that precede the first C label in an
'entry block. The panic! after the block is there for the Rust compiler: the
function returns from inside the state machine, but the compiler cannot see that
every path does, so without a final diverging statement it rejects the function
for not returning a value.
The macro expands to a loop over a match on a state variable, one arm per
block. Each arm ends by setting the next state and continuing the loop, and
goto!('label) sets the target state instead. In outline, the block above
becomes:
#![allow(unused)]
fn main() {
let mut state: u32 = 0;
'sm: loop {
match state {
0 => {
/* entry body */
state = 1;
continue 'sm;
}
1 => {
/* again body, with goto!('again) as */
{
state = 1;
continue 'sm;
}
break 'sm;
}
_ => break 'sm,
}
}
}
break and continue written inside a block (outside any loop nested in it)
still refer to the loop enclosing the goto_block!: the macro records them in a
flag, leaves the state machine loop, and re-issues them after it. goto!
outside a goto_block! is a compile error.
Supported goto patterns:
- labels at the top level of a block: a function body, a loop body, or a compound statement;
- a
gotoanywhere inside that block, including in nestedifs, loops, andswitchcases; - forward and backward jumps.
Not supported yet:
- a jump to a label that is not at the top level of a block enclosing the
goto, such as from outside a loop to a label in its body; - a jump to a label inside an
ifbranch.
switch
A switch without fallthrough is translated as a plain match inside a labeled
block, where break becomes a break out of that block. When some case falls
into the next, the code generator uses switch! instead. It is written like a
match, but an arm whose body does not end in break continues into the body
of the following arm, as C does:
switch (x) {
case 1:
r += 10;
case 2:
r += 20;
break;
default:
r = -1;
break;
}
#![allow(unused)]
fn main() {
switch!(match (*x.borrow()) {
v if v == 1 => {
(*r.borrow_mut()) += 10;
}
v if v == 2 => {
(*r.borrow_mut()) += 20;
break;
}
_ => {
(*r.borrow_mut()) = -1;
break;
}
});
}
switch! desugars to a goto_block! whose first block dispatches on the
condition to the block of the matching case; the case bodies follow as
consecutive blocks, so falling off the end of one enters the next, and break
leaves the whole switch!. A continue in a case is not captured by the
switch!: as in C, it continues the loop enclosing the switch, and is a
compile error when there is none. goto and switch mix freely: a switch!
can be nested in a goto_block!, a goto! inside a case can target a label of
the enclosing block, and a label attached to a case is supported. Statements
between the switch and its first case are not supported yet.
Hoisted declarations
In C a variable declared in one case is visible in the cases after it, because
they all belong to the same block. Each switch! arm is a separate Rust block,
so the code generator hoists such declarations above the macro and leaves an
assignment in the case:
switch (x) {
case 1:
r = 1;
int y;
y = 10;
r += y;
case 2:
y = 20;
r = y;
break;
}
#![allow(unused)]
fn main() {
let y: Value<i32> = <Value<i32>>::default();
switch!(match (*x.borrow()) {
v if v == 1 => {
(*r.borrow_mut()) = 1;
*y.borrow_mut() = 10;
(*r.borrow_mut()) += *y.borrow();
}
v if v == 2 => {
*y.borrow_mut() = 20;
(*r.borrow_mut()) = *y.borrow();
break;
}
_ => {}
});
}
The same hoisting applies to variables used across the labeled blocks of a
goto_block!, as count and acc above show.
I/O and Formatting
The io, format, and fd modules support the stdio stream functions,
printf-style formatting, and descriptor-based I/O.
C Streams
A C FILE is more than a file handle: it carries sticky end-of-file and error
flags that feof and ferror report long after the read that set them.
std::fs::File keeps no such state, so the refcount model translates FILE *
as a Ptr<CFile>, a libc shim that holds the file descriptor
together with these two flags. The standard streams are thread-local CFile
values over descriptors 0, 1, and 2, returned by c_stdin, c_stdout, and
c_stderr. CFile does no buffering at present: every read or write on it is a
system call on the descriptor.
In the unsafe model streams stay raw: stdin_unsafe, stdout_unsafe, and
stderr_unsafe return the process’s *mut libc::FILE handles, whose symbol
names differ per platform (stdin on Linux, __stdinp on macOS).
fread and fwrite exist in both models as named functions, because translated
programs take their address:
#![allow(unused)]
fn main() {
pub fn fread_refcount(
a0: AnyPtr,
a1: usize,
a2: usize,
a3: Ptr<CFile>,
) -> usize;
pub unsafe fn fread_unsafe(
a0: *mut c_void,
a1: usize,
a2: usize,
a3: *mut libc::FILE,
) -> usize;
}
The refcount variant reinterprets the destination as a byte array and reads
through the CFile; the unsafe variant forwards to libc::fread.
C++ Streams
In the refcount model cin, cout, and cerr are translated as
Ptr<std::fs::File> values over duplicates of the standard descriptors,
returned by the cin, cout, and cerr functions. Ptr<T> implements
write_fmt and write_all whenever T: Write, forwarding to the pointee
through with_mut, so cout << x becomes write!(cout(), "{}", x) and a raw
byte range is written with cout().write_all(..). In the unsafe model
cin_unsafe, cout_unsafe, and cerr_unsafe return raw pointers to
thread-local std::fs::File values. C++ streams do not map fully onto
std::fs::File, so this translation may change in the future.
Formatting
The code generator first translates the printf family into the idiomatic
print! and println! macros. That is not always possible: the target stream
may not be known at translation time, or the format string may be a runtime
value. For those cases, and for functions that format into a buffer such as
snprintf, the refcount model falls back to format_c; the unsafe model calls
libc directly.
format_c evaluates a C format string against a slice of
variadic arguments and returns the formatted String:
#![allow(unused)]
fn main() {
pub fn format_c(fmt: &str, va: &[VaArg]) -> String;
}
Parsing and rendering come from the sprintf crate. The integer, character,
string, and floating-point conversions are supported, and %s reads the
argument through the refcounted pointer as a Rust string. A malformed format
string or an argument of the wrong kind is a panic. Three things are not
supported yet:
%prenders through the pointer’s integer cast, which currently panics.%nis not handled.- A
*width or precision (%*d,%.*s) is parsed but its integer argument is not consumed, so the remaining arguments are misaligned.
File descriptors
Rust tracks descriptor ownership in the type system: an OwnedFd closes the
descriptor when dropped, and a BorrowedFd grants temporary access to one. C
has no such distinction: a descriptor is a plain int, mixed freely with
integer arithmetic, so the translator cannot tell which int values are
descriptors. The refcount model therefore leaves descriptors as integers in the
translated program and keeps the ownership in one place, the thread-local
FdRegistry, a table from each integer to the open descriptor it names.
The registry follows the descriptor’s life. When a rule opens a file,
FdRegistry::register stores the resulting OwnedFd and hands the program its
raw number. When a rule performs I/O on that number, FdRegistry::with_fd looks
the entry up and lends it out as a BorrowedFd for the duration of the call.
When the program calls close, FdRegistry::close removes the entry, which
closes the descriptor. The registry starts out holding the standard descriptors
0, 1, and 2.
In the fstat rule, the descriptor argument goes through with_fd:
#![allow(unused)]
fn main() {
fn f2(a0: i32, a1: Ptr<Stat>) -> i32 {
match FdRegistry::with_fd(a0, |fd: BorrowedFd<'_>| {
nix::sys::stat::fstat(fd)
}) {
// ...
}
}
}
with_fds borrows several descriptors at once for select-style calls. The
select rule collects every descriptor set in the fd_set arguments and
borrows them all for the duration of the call:
#![allow(unused)]
fn main() {
let wanted: Vec<i32> = /* the descriptors set in the fd_set arguments */;
FdRegistry::with_fds(&wanted, |borrowed: &[BorrowedFd<'_>]| {
let mut read_set = nix::sys::select::FdSet::new();
for fd in &borrowed[..read_count] {
read_set.insert(*fd);
}
// ... build the write and except sets the same way ...
nix::sys::select::select(nfds, &mut read_set, /* ... */)
})
}
Using a descriptor that was never opened, or using it after it was closed, is a
bug in the original program. The registry turns such a use into a panic (with a
message prefixed ub:) so the bug surfaces instead of going unnoticed.
libc Shims
libc structs hold raw pointers, which are incompatible with the refcounted
pointers the refcount model uses, so a libc struct cannot be used directly. The
shim modules therefore define Rust counterparts for the libc types translated
programs use. A shim struct mirrors its C struct member by member, like a
translated struct: fields are stored inline, except for arrays, which are
Value<Box<[T]>>s of their own (see Boxing). A
shim converts to or from the underlying libc or nix type at the call boundary.
Stat is a typical shim:
#![allow(unused)]
fn main() {
#[derive(Clone, Default, Record)]
pub struct Stat {
#[offset(offset_of!(::libc::stat, st_dev))]
pub st_dev: u64,
#[offset(offset_of!(::libc::stat, st_ino))]
pub st_ino: u64,
// ...
#[offset(offset_of!(::libc::stat, st_size))]
pub st_size: i64,
}
impl Stat {
pub fn from_libc(s: &::libc::stat) -> Self { /* ... */ }
}
}
A stat call in the source program becomes a nix::sys::stat::stat call. On
success nix returns a raw libc::stat, so the result goes through
Stat::from_libc before it is written into the translated struct.
Like translated structs, shims derive Record, so that pointers to their fields
can be taken (see Pointers to fields). The offsets
of the fields are those of the libc struct, given by offset_of!, and their
ByteRepr gives the size of the libc struct, which locates the fields of the
elements of an array of shims, like an array of pollfd. The sockaddr family,
whose byte representation has a layout of its own, uses the offsets of that
layout.
The modules
Each shim lives next to the rules that use it, as rules/<dir>/shim.rs. The
libcc2rs build script finds every such file and includes it as a module of the
crate, re-exported at the crate root, so a shim refers to other runtime items
through crate:: and translated code reaches it as libcc2rs::Stat.
| Rule dir | C types |
|---|---|
stdio | FILE (CFile) |
dirent | struct dirent, DIR (Dirent, CDir) |
select | fd_set (CFdSet) |
ifaddrs | struct ifaddrs (Ifaddrs) |
ip | struct in_addr, struct in6_addr (InAddr, In6Addr) |
netdb | struct addrinfo (Addrinfo) |
poll | struct pollfd (Pollfd) |
pwd | struct passwd (Passwd) |
socket | the sockaddr family (Sockaddr, SockaddrIn, SockaddrIn6, SockaddrUn, SockaddrStorage) |
stat | struct stat (Stat) |
termios | struct termios, struct winsize (Termios, Winsize) |
time | struct tm, struct timeval, struct timespec (Tm, Timeval, Timespec) |
Most shims are plain data plus conversions like Stat. CFile carries the
stdio stream logic (see I/O and Formatting), and the time shims
convert through the jiff crate. CFdSet and the sockaddr family need more
than a field-by-field mirror and are described in their own sections below.
Each shim file also gives the raw libc struct it mirrors an empty ByteRepr
impl (impl ByteRepr for ::libc::stat {}), whose methods panic. These exist so
that the generated ByteRepr implementation of a translated struct with a libc
struct member still compiles; reinterpreting such a struct is not supported at
present.
CFdSet
nix has its own FdSet, but it is stricter than the C one: it ties the set to
the lifetimes of the descriptors it holds. A C fd_set is just a set of
integers that accepts anything; whether the descriptors are valid is only
checked by the select call that eventually receives the set. CFdSet keeps
the C behavior by storing plain integers, and the select rule builds the nix
FdSet from it at call time.
The sockaddr family
C socket code reinterprets one address struct as another: the program fills in a
struct sockaddr_in, passes it to bind as a struct sockaddr *, and casts
back to the concrete type on the way out of accept. The address shims keep
this pattern working by implementing ByteRepr with the
exact byte layout of their C structs: the family in the first two bytes, the
remaining members at their C offsets. A cast in the source program becomes a
reinterpret_cast on the refcounted pointer, which reads
the struct through that byte layout as the target type, so any member of the
family can be viewed as any other, exactly as in C.
The call boundary works the same way. Sockaddr::decode reads the family from
the first two bytes and reinterprets the pointer as the concrete type before
handing nix a typed address:
#![allow(unused)]
fn main() {
pub fn decode(
addr: &Ptr<Sockaddr>,
_len: u32,
) -> Option<Box<dyn SockaddrLike>> {
let family = addr.reinterpret_cast::<u16>().read();
if family == libc::AF_INET as u16 {
let m = addr.reinterpret_cast::<SockaddrIn>().read();
Some(Box::new(nix::sys::socket::SockaddrIn::from(m.to_libc())))
}
// ... AF_INET6 and AF_UNIX in the same way ...
}
}
Sockaddr::encode goes the other way, writing an address returned by nix into
the caller’s buffer through the concrete shim. Ifaddrs hands out its addresses
as Ptr<Sockaddr> values ready to be reinterpreted.
Non-uniform fields
Some struct fields are not spelled the same on every platform. struct stat
keeps the modification time in a nested struct timespec, named st_mtim on
Linux and st_mtimespec on macOS, while the shim exposes a single st_mtime
field. struct in6_addr hides its bytes behind the internal __in6_u union on
Linux, while the shim exposes s6_addr. The shims pick one uniform field, and
the code generator meets them halfway: replaceNonUniformLibcField in the
converter rewrites the platform-specific member chain in the source, so
st.st_mtim.tv_sec becomes st.st_mtime in the translated code.
Compat Helpers
Some C interfaces are macros or platform-specific symbols rather than plain
functions. On the source side, cpp2rust rewrites them into ordinary calls (see
Compat Shims); the compat module is the runtime side of
that rewrite.
errno expands to a platform-specific function call (__errno_location on
Linux, __error on macOS).
In the unsafe model, cpp2rust_errno_unsafe binds both platform symbols under
one name and returns the real libc errno location:
#![allow(unused)]
fn main() {
pub unsafe fn cpp2rust_errno_unsafe() -> *mut i32;
}
In the refcount model, errno is a thread-local refcounted i32 that the
runtime maintains itself:
#![allow(unused)]
fn main() {
pub fn cpp2rust_errno() -> Ptr<i32>;
}
Refcount code reaches the operating system through the libc shims and nix, so
libc’s errno is never read by this model. Keeping the cell current is a
discipline of the rules: every rule that translates a call that can fail must
write the error code into cpp2rust_errno() on the failure path (see
Compat Shims); nothing enforces this, and a rule that
skips the write breaks programs that check errno.
malloc_usable_size is bound under one name for both platforms (the symbol is
malloc_size on macOS).
Overview
This part of the book documents the internals of the code generator: how the clang AST is traversed and how Rust code is emitted.
The Translation Pipeline
flowchart TD
driver["<b>cpp2rust</b><br/><code>cpp2rust/cpp2rust.cpp</code>"]
lib["<b>TranspileSrc / TranspileDir</b><br/><code>cpp2rust/cpp2rust_lib.cpp</code>"]
action["<b>FrontendAction, ASTConsumer</b><br/><code>cpp2rust/ast_consumer.cpp</code>"]
factory["<b>CreateConverter</b><br/><code>cpp2rust/converter/factory.cpp</code>"]
mapper["<b>Mapper::LoadTranslationRules</b><br/><code>cpp2rust/converter/mapper.cpp</code>"]
conv["<b>Converter / ConverterRefCount</b><br/><code>cpp2rust/converter/</code>"]
out["<b>output file</b><br/>rustfmt"]
driver -->|"source or compilation database"| lib
lib -->|"one per translation unit"| action
action --> factory
factory -.->|"first call only"| mapper
factory --> conv
conv -->|"rs_code"| out
The stages
cpp2rust(cpp2rust/cpp2rust.cpp) parses the flags, resolves the rules directory (see Loading and Matching), and callsTranspileSrcfor--fileorTranspileDirfor--dir.TranspileSrc/TranspileDir(cpp2rust/cpp2rust_lib.cpp) run clang tooling over the source or over every file incompile_commands.json, with oneFrontendActionper translation unit.ASTConsumer::HandleTranslationUnitcallsCreateConverter(cpp2rust/converter/factory.cpp), which loads the translation rules on its first call and constructs aConverter(--model=unsafe) or aConverterRefCount(--model=refcount).- The converter emits the file preamble if this is the first unit, then
traverses the unit and appends Rust text to
rs_code. - The driver writes
rs_codeto the-opath and runsrustfmton it.
Gotchas
- Every unit is parsed with the platform flags from
cpp2rust/compat/platform_flags.h, which put the compat headers ahead of the system headers and set-D_FORTIFY_SOURCE=0, so macro-heavy libc APIs reach the converter as plain function calls. - In
--dirmode__FILE__is redefined to the file’s basename, so the generated code does not embed the absolute paths of the build machine. - Rules are loaded once per process, and the file preamble is emitted once, by
the first unit; the bookkeeping that spans units (which declarations and
records have already been emitted) is kept in
staticmembers ofConverter. - After the last unit,
Converter::EmitOpaqueRecordsappendspub struct Name;for every record type that was referenced but never defined, so types only used behind pointers still compile. - A failing
rustfmtis reported as an error, but the unformatted file stays on disk for inspection.
Types
Every place the converter prints a type goes through Convert(QualType). It
first asks the type rules for a mapping, so library
types and typedef names such as size_t are resolved by rules, and only falls
back to the Visit*Type methods for the built-in and user-defined types
described here.
Given
struct Item {
int id;
char name[8];
std::vector<int> refs;
};
int count(Item item) { return item.id; }
the unsafe model produces (attributes and trait impls omitted)
#![allow(unused)]
fn main() {
pub struct Item {
pub id: i32,
pub name: [libc::c_char; 8],
pub refs: Vec<i32>,
}
pub unsafe fn count_0(mut item: Item) -> i32 {
return item.id;
}
}
and the refcount model produces
#![allow(unused)]
fn main() {
pub struct Item {
#[offset(0)]
pub id: i32,
#[offset(4)]
pub name: Value<Box<[u8]>>,
#[offset(16)]
pub refs: Value<Vec<i32>>,
}
pub fn count_0(item: Item) -> i32 {
let item: Value<Item> = Rc::new(RefCell::new(item));
return (*item.borrow()).id;
}
}
Fields are stored inline in their struct, so a whole struct lives in a single
Value, like the elements of an array; only arrays and vectors are Values of
their own (see Boxing). A pointer to a field records the
allocation of the struct plus the byte offset of the field in it, which the
#[offset(N)] attributes give (see
Pointers to fields).
Type Mappings
The table gives the spelling of each C++ type in both models, before any
refcount boxing. T stands for the translated inner type.
| C++ | Unsafe model | Refcount model |
|---|---|---|
bool | bool | bool |
int, unsigned long, … | i32, u64, … (host width) | same |
float, double | f32, f64 | same |
char | libc::c_char | u8 |
size_t and other typedefs | by type rule (usize), else desugared | same |
T[N] | [T; N] | Box<[T]> |
T[] | [T] | Box<[T]> |
struct S, enum E | S, E | same |
T *, T & | *mut T, *const T | Ptr<T> |
Abstract * | *mut dyn Abstract | PtrDyn<dyn Abstract> |
void * | *mut ::libc::c_void | AnyPtr |
R (*)(A) | Option<unsafe fn(A) -> R> | FnPtr<fn(A) -> R> |
va_list | VaList | VaList |
| lambda closure | impl Fn(A) -> R as a parameter, _ elsewhere | same |
std::unique_ptr<T> | by type rule (Option<Box<T>>) | by type rule (Option<Value<T>>) |
std::vector<T> and other STL | by type rule (Vec<T>) | by type rule (Vec<T>, Vec<Value<Vec<T>>> when nested) |
Other built-ins (wchar_t, long double, char16_t) are omitted. Rvalue
references (T &&) have no mapping of their own; they reach the converter only
through std::move and implicit move constructors, which are handled by rules
and by the constructor translation.
User-defined types as rules
When a record or enum declaration is converted,
Mapper::AddRuleForUserDefinedType registers it in the mapper’s type table: the
C++ name maps to the Rust name, and its pointer form maps to *mut Name or
Ptr<Name> (*mut dyn Name or PtrDyn<dyn Name> for abstract classes); nested
records are registered too. This is what makes library types instantiated with
user types translatable: Mapper::Map matches std::vector<Item> against the
rule for std::vector<T1> and then has to map T1 = Item through the same
table, which would fail if Item were not in it.
Scalars
char is libc::c_char in the unsafe model, whose signedness follows the
platform like C’s, and u8 in the refcount model, because C strings are byte
vectors there (see C Strings). Since most C
implementations have signed char, the refcount model is set to switch to i8
(#246).
Arrays
In the refcount model a constant array becomes Box<[T]>, dropping the length.
Ptr<T> carries only the element type, not N, so a [T; N] could not be
pointed to without a Ptr per length; Box<[T]> gives arrays of every length,
and heap arrays, the same shape. [T; N] survives only inside sizeof, which
becomes ::std::mem::size_of::<[T; N]>().
Array parameters decay to pointers as in C.
Typedefs and qualifiers
Typedef names are looked up as type rules before being desugared, which is how
size_t maps to usize instead of the underlying unsigned long.
Constness is dropped: in the unsafe model it survives only as *const on
pointers and as a missing mut on bindings, and in the refcount model it has no
representation.
The Pointers and References page covers how values of pointer types are read and written.
Boxing
In the refcount model a variable is boxed: its type T is wrapped in
Value<T>, an alias for Rc<RefCell<T>> (see
Reference Counting). Without the box, taking the address
of a variable would need a Rust reference, and arbitrary C++ aliasing cannot be
expressed with references.
Not every type position is boxed. ConverterRefCount keeps a stack of
conversion kinds, conversion_kind_, and the construct
that owns the type pushes one before printing it:
FullRefCount: pushed by variable declarations;Convert(QualType)wraps the result inValue<...>.Pointee: pushed by field declarations; the bare type is printed. Fields that are arrays, or whose type maps to aVecor aBox(std::vector,std::string,std::array), pushFullRefCountinstead (see below).Unboxed: pushed by parameter lists, return types, and record names; the bare type is printed.Ptr: pushed by a pointer type for its pointee; also printed bare.
The result by position:
| Position | int | Item | int[3] |
|---|---|---|---|
| local variable, global | Value<i32> | Value<Item> | Value<Box<[i32]>> |
| function parameter, return type | i32 | Item | decays to Ptr<i32> |
| struct field | i32 | Item | Value<Box<[i32]>> |
pointee of Ptr<T>, element of a container | i32 | Item | Box<[i32]> |
Parameters arrive unboxed and are re-boxed by the function preamble; return values are unboxed:
int add(int a, Item item) { return a + item.id; }
#![allow(unused)]
fn main() {
pub fn add_0(a: i32, item: Item) -> i32 {
let a: Value<i32> = Rc::new(RefCell::new(a));
let item: Value<Item> = Rc::new(RefCell::new(item));
return *a.borrow() + (*item.borrow()).id;
}
}
C++ passes arguments to functions by copy, so signatures stay unboxed; boxing
the copy on entry then lets the body treat parameters exactly like local
variables. The preamble skips reference parameters, which are a Ptr<T> and
never boxed.
Nested containers, library ones and arrays alike, box each level except the
innermost, so that every inner container can be borrowed and mutated on its own,
and a pointer can be taken to it. The boxing is written into the type rules
themselves: std::vector<std::vector<int>> maps to Vec<Value<Vec<i32>>>, and
the carray rules map int a[2][2] to Box<[Value<Box<[i32]>>]>, both before
the outer Value<...> of the declaration is added.
Struct fields are stored inline, so that a whole struct is a single allocation,
and a pointer to a field records the struct’s allocation and the field’s byte
offset (see Pointers). Arrays and vectors are the exception: an
array field is a Value<Box<[T]>> of its own, and a vector field a
Value<Vec<T>>. A pointer to an element, or to a field of an element, then has
the array or the vector as its allocation instead of the struct, and pointer
arithmetic moves between elements as for any other array:
struct Holder { std::vector<Point> points; int n; };
h.points[0].y = 5;
#![allow(unused)]
fn main() {
pub struct Holder {
#[offset(0)]
pub points: Value<Vec<Point>>,
#[offset(24)]
pub n: i32,
}
(*(*h.borrow()).points.borrow_mut())[(0_usize) as usize].y = 5;
}
An array or vector field is accessed like a local one, through its own
borrow() or borrow_mut(), and the struct is only borrowed immutably to reach
it. As Value is shared on clone(), structs with such fields implement
Clone by copying the arrays and vectors, instead of deriving it.
Naming
Rust has one flat namespace per module and no overloading, so C++ names are flattened and disambiguated when they are emitted.
Records and enums are named by Mapper::ToRustName from their qualified C++
spelling: ::, <, >, commas, and spaces all become _. So ns::Foo is
ns_Foo, the instantiation MyContainer<int> is MyContainer_int_, and a
struct Level1 nested in Level0 is Level0_Level1. The same name is used for
the struct, its impl blocks, and every mention of the type.
An anonymous struct, union, or enum is named anon_N, numbered in order of
first appearance. In C, an anonymous tag that is only reachable through a
typedef (typedef struct { ... } Point;) is emitted as Point_struct (or
Point_enum), because C keeps tags and ordinary identifiers in separate
namespaces and Point may already be a variable or function.
Names that are Rust keywords get a trailing underscore: a variable type
becomes type_. The same applies to a keyword followed only by underscores, so
a C++ identifier that was already type_ becomes type__ and cannot collide
with the renamed type.
Free functions and global variables get a numeric suffix (main_0, foo_3)
from a process-wide table keyed by mangled name, which keeps overloads and
same-named static functions from different files apart. Methods keep their
name unless they are overloaded, in which case the parameter types are appended
(method_i32, method_i32_const). operator< is emitted as lt; comparison
operators additionally produce the corresponding trait impls (PartialOrd,
Ord, PartialEq).
Copy and move constructors are named copy_from and move_from, and copy and
move assignment operators copy_assign and move_assign, as long as the class
has only one member of that kind and no method already uses the name; otherwise
they get the overloaded form.
Classes and Structs
A class becomes a struct with one field per data member, an impl block holding
its constructors and methods, and trait implementations after it. Given
class Counter {
int count_;
public:
Counter(int start) : count_(start) {}
~Counter() { count_ = 0; }
int get() const { return count_; }
void set(int v) { count_ = v; }
};
the unsafe model produces
#![allow(unused)]
fn main() {
#[repr(C)]
#[derive(Copy, Clone, Default)]
pub struct Counter {
count_: i32,
}
impl Counter {
pub unsafe fn Counter(mut start: i32) -> Self {
let mut this = Self { count_: start };
this
}
pub unsafe fn get(&self) -> i32 {
return self.count_;
}
pub unsafe fn set(&mut self, mut v: i32) {
self.count_ = v;
}
}
}
and the refcount model produces
#![allow(unused)]
fn main() {
#[derive(Default)]
pub struct Counter {
count_: Value<i32>,
}
impl Counter {
pub fn Counter(start: i32) -> Self {
let start: Value<i32> = Rc::new(RefCell::new(start));
let mut this = Self {
count_: Rc::new(RefCell::new(*start.borrow())),
};
this
}
pub fn get(&self) -> i32 {
return *self.count_.borrow();
}
pub fn set(&self, v: i32) {
let v: Value<i32> = Rc::new(RefCell::new(v));
*self.count_.borrow_mut() = *v.borrow();
}
}
impl Drop for Counter {
fn drop(&mut self) {
*self.count_.borrow_mut() = 0;
}
}
impl Clone for Counter {
fn clone(&self) -> Self {
let mut this = Self {
count_: Rc::new(RefCell::new(*self.count_.borrow())),
};
this
}
}
impl ByteRepr for Counter { /* byte_size, to_bytes, from_bytes */ }
}
The unsafe model adds #[repr(C)] and derives what it can; the refcount model
writes most impls by hand. Which traits are emitted, and when they are derived
rather than written, is on the Traits page.
Fields keep their C++ access: pub for public members, nothing for private
ones. A class nested in another class is emitted as its own top-level struct,
named Outer_Inner (see Naming); Rust has no nested types, and
the outer struct refers to it by that name. A record that is only
forward-declared, or whose definition is never converted because it is only used
behind pointers, is emitted at the end of the file as an empty
pub struct Name; (see The Translation Pipeline).
A constructor becomes an associated function named after the class. It opens
with let mut this = Self { ... }, one field per member initializer, then runs
the C++ body and returns this. Copy and move constructors become copy_from
and move_from (see Naming); the Clone impl calls copy_from
when the copy constructor is user-defined.
Methods take &self when const and &mut self otherwise in the unsafe model.
In the refcount model they always take &self, since mutation goes through the
fields’ RefCells. Inside a method, this is self.
A destructor with a body becomes impl Drop in the refcount model.
Warning
The unsafe model does not emit destructors at all; a user-defined destructor is silently dropped (#310).
Inheritance
An abstract class becomes a trait with one method per pure virtual function, and a class deriving from it implements the trait with its overrides. Given
class Animal {
public:
virtual bool bark() const = 0;
};
class Dog : public Animal {
bool bark() const override { return true; }
};
the unsafe model produces (attributes omitted)
#![allow(unused)]
fn main() {
pub unsafe trait Animal {
unsafe fn bark(&self) -> bool;
}
pub struct Dog {}
unsafe impl Animal for Dog {
unsafe fn bark(&self) -> bool {
return true;
}
}
}
and the refcount model produces (attributes and the Clone and ByteRepr impls
omitted)
#![allow(unused)]
fn main() {
pub trait Animal {
fn bark(&self) -> bool;
}
pub struct Dog {}
impl Animal for Dog {
fn bark(&self) -> bool {
return true;
}
}
}
Non-virtual methods of the derived class go into its own impl Dog block as
usual. Because the base is a trait, pointers to it are *mut dyn Animal in the
unsafe model and PtrDyn<dyn Animal> in the
refcount model, and a Dog * is upcast at the call site. Only the first base
class is considered, and only virtual methods go through the trait; bases with
data members or non-virtual methods, and multiple inheritance, are outside the
supported subset.
Templates
Class templates are translated by full instantiation: each instantiation used by
the program becomes its own struct and impl block, named after the template
arguments (see Naming). MyContainer<int> and
MyContainer<char> become MyContainer_int_ and MyContainer_char_, each with
a complete copy of the methods specialized for its element type. Nothing is
shared between instantiations, and Rust generics are not used.
Flexible array members
A trailing array member of size 0, 1, or [] that C code over-indexes into
memory allocated past the struct is detected with clang’s
isFlexibleArrayMemberLike. In the unsafe model an access to such a member is
not an array index, which Rust would bounds-check against the declared length,
but pointer arithmetic from the array’s start: s.bytes[i] becomes
*s.bytes.as_mut_ptr().add(i as usize), and &s.bytes[i] the same without the
leading *. The refcount model has no dedicated handling; the pattern works
when the array is a union member, because the union accessor returns a Ptr
over the whole allocation that can be offset freely.
Traits
Every emitted struct comes with a fixed set of trait implementations. Some are
derived, some are written out; which is which depends on the model. Enums derive
Clone, Copy, PartialEq, Debug, Default in both models, and unions derive
Copy, Clone in the unsafe model regardless of their fields; the table below is
for structs.
| Trait | Unsafe model | Refcount model |
|---|---|---|
Copy | derived when every field is copyable | never |
Clone | derived | derived when member-wise, else hand-written |
Default | derived when possible, else hand-written | same |
Drop | not emitted (#310) | hand-written from a user destructor with a body |
Ord, PartialOrd, PartialEq, Eq | hand-written from operator< | same |
ByteRepr | not needed | hand-written for every record and enum |
Record | not needed | derived for every struct |
Copy and Clone
In the unsafe model a record derives Copy unless a field is translated to a
Vec, BTreeMap, Option<Box<T>> (std::unique_ptr), or a record that is not
itself Copy; Clone is derived unless the C++ copy constructor is deleted.
The refcount model skips Clone altogether when the copy constructor is
deleted, and never derives Copy. It derives Clone for C structs and for
classes with an implicit or defaulted copy constructor, which copy each field
with its own clone, i.e., its C++ copy constructor. The exceptions are fields
that are or nest a Value, such as std::vector<int> (Value<Vec<i32>>, see
Boxing) or std::vector<std::vector<int>>
(Value<Vec<Value<Vec<i32>>>>), whose derived clone would share the Values
instead of copying them; the generated Clone then translates the implicit copy
constructor, which copies them deeply.
A user-defined copy constructor takes a Ptr to the source, while clone only
has &self. clone passes it a pointer to a shallow copy of self, built
field by field without running any copy constructor, so that the source is
copied exactly once.
This is what gives struct assignment and pass-by-value C++’s member-by-member copy.
Default
Default is the value of a T x; without initializer, of T x = {}, and of
the elements of new T[n]. It is derived when the derived impl gives the C zero
value, and hand-written otherwise: when the class has a user-defined default
constructor, default() calls it; when a field is a C array, a std::array, a
function pointer, or a libc record, default() builds the struct field by
field, each with the same default value the converter uses for a variable of
that type declared without an initializer:
#![allow(unused)]
fn main() {
impl Default for S {
fn default() -> Self {
S {
head: 0_i32,
tail: [0_i32; 3],
buf: [0 as libc::c_char; 4],
}
}
}
}
Unions always get a hand-written impl that zeroes their bytes.
Drop
A user-defined destructor with a non-empty body becomes impl Drop, with the
body translated as a method body. Only the refcount model emits it; the unsafe
model drops destructors silently
(#310).
Comparison
A class that defines operator< (as a method or an out-of-line function) gets
Ord, PartialOrd, PartialEq, and Eq, all expressed through the emitted
lt method: cmp calls it both ways to pick Less, Greater, or Equal, and
eq is “neither is less”. Only one comparison operator per class is supported,
and only operator<. The converter assumes the operator is const, which Rust
requires (cmp and eq take &self) but C++ does not; a non-const
operator< is still emitted as lt(&self, ...).
ByteRepr
The refcount model emits ByteRepr for
every record and enum: byte_size, to_bytes, and from_bytes laid out with
the C offsets of the fields (enums go through their i32 value). It is what
lets a Ptr to the type be reinterpreted as bytes, and bytes be read back as
the type. A record with a field that has no byte representation gets an impl
with byte_size alone, which locates the elements of arrays of the record for
field pointers, and reinterpreting it panics at run time.
Record
#[derive(Record)], from libcc2rs-macros, gives
pointers to fields access to the
fields of a struct. The converter writes the C offset of each field, from
Clang’s record layout, as an #[offset(N)] attribute on it, and the derive
generates the table of offsets that field_ptr! looks fields up in, and the
locate functions that find a field from its offset when a field pointer is
accessed.
Enums
An enum becomes a Rust enum with one variant per enumerator and explicit
discriminants, followed by the impls that give it C semantics. Given
enum Color { RED, GREEN, BLUE };
both models produce
#![allow(unused)]
fn main() {
#[derive(Clone, Copy, PartialEq, Debug, Default)]
enum Color {
#[default]
RED = 0,
GREEN = 1,
BLUE = 2,
}
impl From<i32> for Color {
fn from(n: i32) -> Color {
match n {
0 => Color::RED,
1 => Color::GREEN,
2 => Color::BLUE,
_ => panic!("invalid Color value: {}", n),
}
}
}
libcc2rs::impl_enum_inc_dec!(Color);
}
and the refcount model adds an impl ByteRepr for Color that converts through
the i32 value.
The first enumerator is the #[default], which is what a zero-initialized or
default-constructed enum variable holds. From<i32> is the integer-to-enum
cast; a value that matches no enumerator panics, where C would silently keep the
integer. impl_enum_inc_dec! implements the four
++/-- forms by stepping through the enumerators. Enum-to-integer casts need
no impl and become as i32.
enum class is translated the same way; enumerators are always spelled
Color::RED on the Rust side, scoped or not. In C, an anonymous enum named only
through a typedef (typedef enum { ... } Tag;) is emitted as Tag_enum, and an
anonymous enum with no name at all as anon_N (see Naming).
Unions
Given
union Number {
int i;
float f;
};
int foo(void) {
union Number u;
u.i = 42;
return u.i;
}
the unsafe model produces
#![allow(unused)]
fn main() {
#[repr(C)]
#[derive(Copy, Clone)]
pub union Number {
pub i: i32,
pub f: f32,
}
impl Default for Number {
fn default() -> Self {
unsafe { std::mem::zeroed() }
}
}
pub unsafe fn foo_0() -> i32 {
let mut u: Number = <Number>::default();
u.i = 42;
return u.i;
}
}
and the refcount model produces
#![allow(unused)]
fn main() {
pub struct Number {
__bytes: Value<Box<[u8]>>,
}
impl Number {
pub fn i(&self) -> Ptr<i32> {
(self.__bytes.as_pointer() as Ptr<u8>).reinterpret_cast()
}
pub fn f(&self) -> Ptr<f32> {
(self.__bytes.as_pointer() as Ptr<u8>).reinterpret_cast()
}
}
impl Default for Number {
fn default() -> Self {
Number {
// 4 is sizeof(union Number): the size of its largest member
__bytes: Rc::new(RefCell::new(Box::from([0u8; 4]))),
}
}
}
pub fn foo_0() -> i32 {
let u: Value<Number> = Rc::new(RefCell::new(<Number>::default()));
(*u.borrow_mut()).i().write(42);
return (*u.borrow()).i().read();
}
}
(Clone and ByteRepr impls omitted.)
Unsafe model
A union is a Rust union with #[repr(C)] and #[derive(Copy, Clone)], one
field per member with the same types as a struct would have. Rust cannot derive
Default for a union, so a hand-written impl zeroes the bytes, which is also
what C’s zero-initialization gives. Members are read and written like struct
fields; the whole function is unsafe, so no extra block is needed.
Refcount model
Rust unions are unusable from safe code, so the refcount model stores the union
as one byte buffer, __bytes: Value<Box<[u8]>>, sized to the largest member,
and emits one accessor method per member. Each accessor takes a pointer to the
buffer and reinterprets it as the member type,
returning a Ptr<T> (a Ptr to the element type for array members). Member
access u.i therefore becomes a call, u.i(), and the read or write goes
through the pointer: .read() and .write(v) for scalars, .upgrade().deref()
for struct members whose fields are then accessed as usual.
Because every member views the same bytes, writing through one member and
reading through another has C’s semantics: the bytes are reinterpreted, not
converted. This is also why the type must implement ByteRepr; the accessor’s
reinterpret_cast needs the member types to have a byte-level representation.
Default fills the buffer with zeros, Clone copies the buffer into a fresh
Value, and ByteRepr copies the buffer in and out. The Rust struct has no
pub fields, so translated code can reach the storage only through the
accessors.
Warning
Accessors are broken on a reinterpreted union. Reading a
Ptr<U>obtained fromreinterpret_castbuilds a temporaryUfrom the bytes withfrom_bytes, sop.upgrade().deref().i()returns a pointer into that temporary’s buffer, which dangles as soon as the statement ends, and a write through it would never reach the original allocation (#311).
Bit-fields
Bit-fields are not implemented. VisitFieldDecl ignores the declared width, so
a field such as unsigned flags : 3; is emitted as a plain u32 field: the
struct compiles, but its layout, size, and the wrap-around of the field’s value
differ from C.
Warning
Code that relies on bit-field layout or width (packing several fields into one word,
sizeofon such a struct, storing a value wider than the field) translates silently to something else (#312).
Pointers and References
The unsafe model keeps C++ pointers as raw pointers and dereferences them
directly. The refcount model replaces every pointer and reference with
Ptr<T>, a weak reference plus an
offset, and every dereference with a short-lived borrow of the pointee. Given
int f(int *q) {
int b = 2;
int *p = &b;
*p = *q;
return b;
}
the unsafe model produces
#![allow(unused)]
fn main() {
pub unsafe fn f_0(mut q: *mut i32) -> i32 {
let mut b: i32 = 2;
let mut p: *mut i32 = &mut b as *mut i32;
*p = *q;
return b;
}
}
and the refcount model produces
#![allow(unused)]
fn main() {
pub fn f_0(q: Ptr<i32>) -> i32 {
let q: Value<Ptr<i32>> = Rc::new(RefCell::new(q));
let b: Value<i32> = Rc::new(RefCell::new(2));
let p: Value<Ptr<i32>> = Rc::new(RefCell::new(b.as_pointer()));
p.borrow().write(q.borrow().read());
return *b.borrow();
}
}
The rest of the page goes through the pointer operations one at a time.
Address-of
Unsafe model: &x becomes &mut x as *mut T, or &x as *const T when the
pointer type is to const. Globals use &raw mut x so no reference to the
static is formed. An array decays with arr.as_mut_ptr(), and the address of
an element is &mut arr[i] as *mut T.
Refcount model: &x becomes x.as_pointer(), which produces a Ptr holding a
weak reference to the variable’s Value. &s.field is field_ptr!(s, field),
and &p->field is field_ptr!(p, field): a pointer to the field of the struct,
made from the Value or the Ptr of the struct (see
Pointers to fields). An array decays
with arr.as_pointer() as Ptr<T>, a Ptr to element 0 of the whole array, and
&arr[i] is that pointer offset by i. An array field decays with
array_field_ptr!(p, arr), and p->arr[i] is read and written through that
pointer offset by i.
Both models push the address down to the innermost place expression:
&(cond ? x : y) becomes if cond { &mut x } else { &mut y } in the unsafe
model and if cond { x.as_pointer() } else { y.as_pointer() } in the refcount
one. The if yields a value copied out of whichever branch ran, not that
branch’s storage, so the address has to be taken inside the branches, where the
place is still known.
Dereference
Unsafe model: *p stays *p, p->x becomes (*p).x, and an assignment
through a pointer is *p = v.
When the dereference is the base of an index, (*p)[i] with p: *mut Vec<i32>,
Rust would have to create a &mut Vec<i32> out of the raw pointer to call
Index::index, and its dangerous_implicit_autorefs lint rejects that as an
error. The converter makes the reference explicit instead: EmitDeref prints
(&mut (*p))[i], or (&(*p))[i] when the operator[] is const, when
autoref_mut_ is set. PushExplicitAutoref sets it
around the base of an overloaded subscript, which also covers a member of the
pointee, (&mut (*hp)).v[i], around the range of a range-for, and around a
rule placeholder marked is_index_base (see Rules IR);
EmitDeref clears it once used, so nested dereferences inside the base are
printed plainly.
Refcount model: a dereference cannot hand out a &T into the pointee, because
nothing would bound the borrow’s lifetime, so a read copies the value out and a
write copies it in. Which form is emitted follows the
expression kind, what the enclosing construct expects
of the dereference:
- An rvalue use copies the value out. A scalar or pointer pointee is
p.read(), and so is a whole record,*p. A field of a record pointee is copied out in a closure that borrows the record for its duration:p->xisp.with(|__s| __s.x), andp->a.bisp.with(|__s| __s.a.b);ReadFieldconverts the record withrecord_ptr_set, so that its dereference is emitted as__s. A field that is aValueof its own, or astd::unique_ptr, is copied out as well, i.e., itsRc:p->v.size()is(*p.with(|__s| __s.v.clone()).borrow()).len(). When the record is not reached through a pointer, the copy is in a block,{ (*s.borrow()).x }. In all cases the record doesn’t stay borrowed for the rest of the statement, which may write to it. - An address-of use prints
pitself. - An lvalue use prints nothing at once. The converter records the pointer
expression as a pending dereference, and
whoever consumes the lvalue, an assignment or a mapped method call, wraps it:
p.write(v), orwith_mutfor a mutating method on a boxed pointee. This is what lets*p = vcome out as a singlewriteinstead of a borrow followed by an assignment, and*p += vas{ let _ptr = p.clone(); _ptr.write(_ptr.read() + v) }. A field of a record pointee is a pending dereference too, offield!(p, x), which projects the pointer to the field (see Pointers to fields):p->x = visfield!(p, x).write(v), andp->a.b = visfield!(field!(p, a), b).write(v).
with and with_mut work on any pointer, including a
reinterpreted one, whose pointee is decoded from
the bytes of the original allocation and, for with_mut, encoded back into them
before the closure returns. A write through a reinterpreted pointer is hence
visible right away through every other pointer to the same bytes.
Arithmetic and comparison
Unsafe model: p + n is p.offset(n as isize) and p - n is
p.offset(-(n as isize)); p - q is
(p as usize - q as usize) / ::std::mem::size_of::<T>(); ++p is
p.prefix_inc() through the increment traits;
p == NULL is p.is_null() and the null literal is std::ptr::null_mut(), or
std::ptr::null() for a pointer to const.
Refcount model: the same operations on Ptr, p.offset(n as isize),
p.clone() - q.clone() (subtraction takes its operands by value, hence the
clones), p.prefix_inc(), p.is_null(), and Ptr::null(). Arithmetic only
moves the offset; whether the result is in bounds is checked when it is
dereferenced.
References
A C++ reference is a pointer that cannot be made to point elsewhere, and both
models translate it as one. In the unsafe model a reference parameter is
*mut T (*const T for const T &), an argument f(x) is
f(&mut x as *mut T), and uses of the reference are *r. In the refcount model
it is a Ptr<T> that is never boxed: the argument is
f(x.as_pointer()), uses are r.read() and r.write(v), and returning a
reference returns the Ptr (with a .clone(), since Ptr is not Copy).
Heap
new T(v) becomes Box::leak(Box::new(v)) as *mut T in the unsafe model and
Ptr::alloc(v) in the refcount model;
delete p becomes ::std::mem::drop(Box::from_raw(p)) and p.delete(). Array
forms use a boxed slice and Ptr::alloc_array. In the refcount model the heap
allocation is a leaked Rc that delete recovers, so a double delete or a
delete of something that was not allocated with new panics instead of
corrupting memory.
Casts
Casts are of two kinds: scalar casts, which both models spell with Rust’s as
or with a small expression, and pointer casts, where the models diverge. Most
casts in the input are implicit, inserted by clang, and are translated the same
way as explicit ones.
Scalar casts
An integer conversion becomes expr as T (an integer literal is instead
re-typed in place: 1 cast to unsigned char prints as 1_u8), and is dropped
when source and target map to the same Rust type, so int to long on a
platform where both are i32 prints nothing. Floating conversions are as as
well. The other scalar casts have their own spellings:
- Integer to
bool:x != 0; a comparison or logical operator that already yieldsboolis left alone. Enum toboolcompares against<E>::from(0). - Pointer to
bool:!p.is_null(). - Integer to enum:
<E>::from(x), theFrom<i32>impl from the Enums page. When the operand is itself a constant of that same enum, which C++ sees as an integer being converted back to the enum, the cast is dropped and the constant is printed directly (Color::RED, not<Color>::from(Color::RED as i32)). Enum to integer isas. - A cast to
void, used to silence an unused-variable warning, becomes a statement that only mentions the operand:&x;in the unsafe model,(*x.borrow()).clone();in the refcount model.
Explicit static_cast, C-style, and reinterpret_cast between scalars follow
the same rules; a cast to the operand’s own type is elided.
Implicit conversions to usize and isize
size_t, size_type, and ssize_t are translated as usize and isize
rather than as the u64/i64 of the unsigned long/long they are typedefs
of (built-in type rules, looked up on the sugared type before it is desugared).
This keeps rules and output free of as usize casts on lengths and indexes, but
it splits one C type in two: clang inserts no conversion between size_t and
unsigned long, while usize and u64 do not mix in Rust. Given
unsigned long take_ulong(unsigned long x);
size_t sz = 20;
unsigned long r = take_ulong(sz);
the refcount model produces
#![allow(unused)]
fn main() {
let sz: Value<usize> = Rc::new(RefCell::new(20_usize));
let r: Value<u64> = Rc::new(RefCell::new(take_ulong_0(*sz.borrow() as u64)));
}
Convert(expr, implicit_convert_to) is the single place where such a cast is
added: the caller passes the type the context expects, NeedsImplicitScalarCast
checks that it is the same C type as the expression’s but maps to a different
Rust type, and if so the expression is wrapped in (...) as <target>. Callers
that pass a target are assignments and initializations (the variable’s type),
call arguments (the parameter type of the callee or rule,
GetParamImplicitConvertTarget, as in the example), and binary operators, which
pick one Rust type for both operands (GetOperandImplicitConversionTarget).
Pointer casts
Given
uint32_t value = 0x04030201;
uint8_t *bytes = (uint8_t *)&value;
void *any = bytes;
uint8_t *back = (uint8_t *)any;
the unsafe model produces
#![allow(unused)]
fn main() {
let mut value: u32 = 67305985_u32;
let mut bytes: *mut u8 = (&mut value as *mut u32) as *mut u8;
let mut any: *mut ::libc::c_void = bytes as *mut ::libc::c_void;
let mut back: *mut u8 = any as *mut u8;
}
and the refcount model produces
#![allow(unused)]
fn main() {
let value: Value<u32> = Rc::new(RefCell::new(67305985_u32));
let bytes: Value<Ptr<u8>> =
Rc::new(RefCell::new(value.as_pointer().reinterpret_cast::<u8>()));
let any: Value<AnyPtr> = Rc::new(RefCell::new((*bytes.borrow()).to_any()));
let back: Value<Ptr<u8>> =
Rc::new(RefCell::new((*any.borrow()).reinterpret_cast::<u8>()));
}
Unsafe model
Every pointer cast, whether written as a C cast, static_cast, or
reinterpret_cast, is a Rust as between raw pointer types. A cast that only
adds or removes const changes the Rust type too, since T * is *mut T and
const T * is *const T, and becomes .cast_const() or .cast_mut(). A cast
that changes nothing in Rust, such as a typedef to its underlying type, is not
emitted. Casts between pointers and integers are also as.
Refcount model
A Ptr<T> is a weak reference to a RefCell<T>, so it cannot simply be
relabeled as a Ptr<U>: the cell it points to holds a T. A cast to another
pointee type therefore produces a different kind of pointer, one that views the
allocation as bytes. Three helpers from the runtime cover the cases:
p.reinterpret_cast::<U>()produces aPtr<U>of theReinterpretedkind: a byte-level view over the original allocation, with the offset counted in bytes. Reads and writes through it go throughByteRepr, which is why every record type gets aByteReprimpl.p.to_any()erases the type into anAnyPtr, the translation ofvoid *, remembering the original type.any.reinterpret_cast::<T>()recovers aPtr<T>from anAnyPtr: the original pointer ifTis the type it was erased from, a byte view otherwise (see AnyPtr casts).
Two casts do not use these helpers. An array decaying to a pointer is spelled
arr.as_pointer() as Ptr<T>, where the as only names the pointer type. An
upcast from a derived class to an abstract base becomes
p.to_dyn::<dyn Base>(|w| w), which is the ordinary Rust unsizing coercion
applied to the pointer’s weak reference (see
Virtual Classes). Casts between pointers and
integers use the integer cast API of Ptr.
Constness is dropped in a cast as everywhere else, so const_cast is a no-op.
dynamic_cast is not supported.
Function pointers
Casting a function pointer to a different signature, which C code does to call
through a generic type, wraps the function in an adapter closure that converts
the arguments; see Casts on the runtime page.
Storing a function pointer in a void * uses to_any() like any other pointer.
Function Pointers
Given
typedef int (*op_t)(int);
int inc(int x) { return x + 1; }
op_t pick() { return inc; }
int apply(op_t f) {
if (f == nullptr) {
return 0;
}
return f(10);
}
the unsafe model produces
#![allow(unused)]
fn main() {
pub unsafe fn inc_0(mut x: i32) -> i32 {
return x + 1;
}
pub unsafe fn pick_1() -> Option<unsafe fn(i32) -> i32> {
return Some(inc_0);
}
pub unsafe fn apply_2(mut f: Option<unsafe fn(i32) -> i32>) -> i32 {
if f.is_none() {
return 0;
}
return f.unwrap()(10);
}
}
and the refcount model produces
#![allow(unused)]
fn main() {
pub fn inc_0(x: i32) -> i32 {
let x: Value<i32> = Rc::new(RefCell::new(x));
return *x.borrow() + 1;
}
pub fn pick_1() -> FnPtr<fn(i32) -> i32> {
return FnPtr::<fn(i32) -> i32>::new(inc_0);
}
pub fn apply_2(f: FnPtr<fn(i32) -> i32>) -> i32 {
let f: Value<FnPtr<fn(i32) -> i32>> = Rc::new(RefCell::new(f));
if (*f.borrow()).is_null() {
return 0;
}
return (*(*f.borrow()))(10);
}
}
Unsafe model
A function pointer is Option<unsafe fn(A) -> R>: Option because it can be
null, unsafe fn because every translated function is unsafe. Naming a
function where a pointer is expected wraps it in Some(...), the null pointer
is None, a null check is is_none(), and a call is f.unwrap()(args).
Function pointers are Copy and compare with == on the function’s address.
Refcount model
A function pointer is FnPtr<fn(A) -> R>, built with
FnPtr::<fn(A) -> R>::new(f). FnPtr dereferences to the function, so a call
is (*f)(args); the null pointer is FnPtr::null() and the check is
is_null(). FnPtr is not Copy, so storing or passing one that already lives
in a variable clones it, and equality compares the address of the wrapped
function, so a pointer stays equal to itself after being cast.
A lambda is also an FnPtr, so a capture-less lambda assigned to a function
pointer is the lambda itself (see Lambdas).
Lambdas
A lambda becomes an FnPtr in both models. Its type
is FnPtr<fn(A) -> R>, with the signature of the lambda’s call operator, so it
can be written wherever C++ names the closure type: a variable, a struct field,
a parameter of an instantiated template, or decltype. A call is
f.call(args).
Given
int total = 0;
auto accumulate = [total](int x) mutable {
total += x;
return total;
};
accumulate(1);
the unsafe model produces
#![allow(unused)]
fn main() {
let mut total: i32 = 0;
let mut accumulate: FnPtr<fn(i32) -> i32> = lambda_unsafe!(
{
let total: i32 = total;
},
|x: i32| -> i32 {
total += x;
return total;
}
);
(unsafe { accumulate.call(1) });
}
and the refcount model produces
#![allow(unused)]
fn main() {
let total: Value<i32> = Rc::new(RefCell::new(0));
let accumulate: Value<FnPtr<fn(i32) -> i32>> = Rc::new(RefCell::new(lambda!(
{
let total: Value<i32> = Rc::new(RefCell::new((*total.borrow())));
},
|x: i32| -> i32 {
let x: Value<i32> = Rc::new(RefCell::new(x));
(*total.borrow_mut()) += (*x.borrow());
return (*total.borrow());
}
)));
({ (*accumulate.borrow()).call(1) });
}
The lambda macros
A lambda with captures is written with lambda! in the refcount model and
lambda_unsafe! in the unsafe model. Both take two arguments:
- a block with one
letper capture, giving its name, its type and the value it is initialized with when the lambda is created; - a closure with the lambda’s parameters, its return type and the translated body.
The macro declares a hidden struct with one field per capture, and makes the
body the call method of that struct. The body is written with the names of the
captures, as the C++ body is, so the macro adds self. in front of each capture
name in it, which makes the name refer to the field of the struct. The lambda!
above expands to
#![allow(unused)]
fn main() {
{
struct __Lambda {
total: Value<i32>,
}
impl __Lambda {
fn call(&self, x: i32) -> i32 {
let x: Value<i32> = Rc::new(RefCell::new(x));
(*self.total.borrow_mut()) += (*x.borrow());
return (*self.total.borrow());
}
}
FnPtr::<fn(i32) -> i32>::from_lambda(
__Lambda {
total: Rc::new(RefCell::new((*total.borrow()))),
},
__Lambda::call,
)
}
}
The struct is local to the expression and its type is erased by the FnPtr, so
it never appears in the translated program. The initializers are not part of the
method: they are evaluated where the lambda is written, so total there is the
enclosing variable, while total in the body is self.total.
The two macros differ in how the method receives the struct. lambda! takes it
by immutable reference, since the captures of the refcount model are Values
and are written through their cells. lambda_unsafe! takes it by mutable
reference, since the captures of the unsafe model are plain fields, and wraps
the body in unsafe.
Every use of a capture name in the body gets the self., so the body cannot
give that name to something else: a variable declared in the body, or a
parameter of a closure in it, with the name of a capture is a compile error.
Captures
The captures are the ones of the C++ closure type, which clang computes also for
[=] and [&].
| Capture | Unsafe model | Refcount model |
|---|---|---|
[x] | T | Value<T> |
[&x] | *mut T | Ptr<T> |
[n = expr] | T | Value<T> |
[this] | *mut S | Value<Ptr<S>> |
[*this] | S | Value<S> |
A capture by copy holds its own copy, made when the lambda is created, and keeps its value between calls. A capture by reference holds a pointer to the variable and is dereferenced in the body:
int base = 10;
auto add_base = [&base](int x) { return x + base; };
#![allow(unused)]
fn main() {
let mut base: i32 = 10;
let mut add_base: FnPtr<fn(i32) -> i32> = lambda_unsafe!(
{
let base: *mut i32 = &mut base;
},
|x: i32| -> i32 {
return ((x) + (*base));
}
);
}
A captured this is the capture this_. Inside the body this refers to it,
and members are accessed through that pointer.
A constant that the body uses without capturing, such as a const int local
read by value, is replaced by its initializer.
Lambdas without captures
A lambda without captures needs no struct and is a function pointer built from a closure:
auto one = [](int x) { return x + 1; };
#![allow(unused)]
fn main() {
let mut one: FnPtr<fn(i32) -> i32> = FnPtr::<fn(i32) -> i32>::new(|x: i32| -> i32 {
unsafe {
return ((x) + (1));
}
});
}
Converting it to a function pointer gives, in the refcount model, the lambda
itself, which already has the type of the function pointer. The unsafe model
translates function pointers as Option<unsafe fn(A) -> R>, so the conversion
emits Some(|...| ...) with the closure again (see
Function Pointers). Lambdas with captures cannot be
converted to function pointers, as in C++.
A lambda without captures can be default-constructed since C++20. The
translation of decltype(one) other; emits the closure again, as do the fields
of that type in the Default of a struct.
Limitations
- Copying a lambda does not copy its captures: the copy shares them with the
original, so a
mutablelambda and its copy update the same state, and the copy constructors of the captures do not run. - Moving a lambda does not run the move constructors of its captures.
- Destroying a lambda does not run the destructors of its captures.
- A lambda that is default-constructed by a type translated with a
rule is wrong. The rule only has the type of
the lambda,
FnPtr<fn(A) -> R>, whose default is the null pointer and not the lambda, so calling it panics. This affects, for example,std::set<int, decltype(cmp)> s;, where the set constructs the comparator. - Generic lambdas, whose call operator is a template, are not translated.
- A default-constructed lambda whose body declares a
staticlocal is rejected, as each emitted closure would get its own copy of the variable. - The parameter and return types of a lambda must implement
FnPtrArg, like those of anyFnPtr.
Special-cased Library Types
Library types are translated by type rules, and for
most of them the converter does nothing beyond applying the rule. A few types
also have code of their own in the converter, which decides how their values are
dereferenced, iterated, or initialized. Some of it could be moved into
rules. This page lists what the converter does
today; the rest of the type comes from its rule module under rules/.
std::unique_ptr
The type rule maps std::unique_ptr<T> to Option<Box<T>> in the unsafe model
and Option<Value<T>> in the refcount model, and std::make_unique and
std::move are rules too. What the converter special-cases (IsUniquePtr in
converter_lib) is everything that treats a unique_ptr as a pointer:
Given
std::unique_ptr<int> x1 = std::make_unique<int>(0);
std::unique_ptr<int> x2 = std::make_unique<int>(0);
*x2 = 1;
x1 = std::move(x2);
int *raw = &*x1;
the unsafe model produces
#![allow(unused)]
fn main() {
let mut x1: Option<Box<i32>> = Some(Box::new(0));
let mut x2: Option<Box<i32>> = Some(Box::new(0));
*x2.as_deref_mut().unwrap() = 1;
x1 = x2;
let mut raw: *mut i32 = &mut (*x1.as_deref_mut().unwrap()) as *mut i32;
}
and the refcount model produces
#![allow(unused)]
fn main() {
let x1: Value<Option<Value<i32>>> =
Rc::new(RefCell::new(Some(Rc::new(RefCell::new(0)))));
let x2: Value<Option<Value<i32>>> =
Rc::new(RefCell::new(Some(Rc::new(RefCell::new(0)))));
*(*x2.borrow_mut()).as_ref().unwrap().borrow_mut() = 1;
*x1.borrow_mut() = (*x2.borrow_mut()).take();
let raw: Value<Ptr<i32>> =
Rc::new(RefCell::new((*x1.borrow()).as_pointer()));
}
*p and p->x are overloaded operator calls in C++; the converter emits the
as_deref_mut().unwrap() and as_ref().unwrap().borrow_mut() forms instead of
an operator call, &*p becomes the raw pointer or as_pointer(), and
std::move of a unique_ptr is a plain move or a take(). In the unsafe model
p == nullptr becomes p.is_none(); the refcount model does not special-case
it and emits is_null() as for any pointer, which no test exercises on an
Option<Value<T>>. A struct with a unique_ptr field does not derive Copy.
Iterators
Iterator types come from rules (std::vector<T>::iterator maps to *mut T or
Ptr<T>, std::map<K, V>::iterator to UnsafeMapIterator<K, V> or
RefcountMapIter<K, V>), and the converter classifies them by
GetStrongestIteratorCategory: rule types marked as refcount pointers are
contiguous iterators and are handled exactly like a Ptr<T>, and the map
iterator types are bidirectional. The classification drives a few decisions.
Given
std::map<int, double> m;
double sum = 0;
for (const auto &i : m) {
sum += i.second;
}
auto it = m.begin();
sum += it->second;
the unsafe model produces
#![allow(unused)]
fn main() {
for i in UnsafeMapIterator::begin(&m as *const BTreeMap<i32, Box<f64>>) {
sum += *i.second();
}
let mut it: UnsafeMapIterator<i32, f64> =
UnsafeMapIterator::begin(&m as *const BTreeMap<i32, Box<f64>>);
sum += *it.second();
}
and the refcount model produces
#![allow(unused)]
fn main() {
for i in RefcountMapIter::begin(m.as_pointer()) {
*sum.borrow_mut() += *i.second().borrow();
}
let it: Value<RefcountMapIter<i32, f64>> =
Rc::new(RefCell::new(RefcountMapIter::begin(m.as_pointer())));
*sum.borrow_mut() += *(*it.borrow()).second().borrow();
}
it->second on a bidirectional iterator (map iterator) is not a pointer
dereference plus a field access, since the map iterator types have no pointer to
hand out; the converter emits the iterator itself and the field rule turns the
access into an accessor call.
The loop variable of a range-for over a std::map is the map iterator itself,
an entry with first()/second() accessors, not a pointer to an element; the
converter remembers such variables in map_iter_decls_
so that uses of them are not dereferenced.
A converting-constructor call that only wraps an iterator does not clone it
(PushSuppressIteratorClone). libstdc++ and libc++ differ in whether such a
wrapping constructor appears in the AST, so skipping the clone keeps the output
identical on Linux and macOS.
IsIteratorType recognizes any record that declares an iterator_category
typedef.
std::array
std::array<T, N> maps to Vec<T> (see rules/array), so an initializer
{1, 2, 3} becomes vec![1, 2, 3]. The converter knows the type by name in
three places: the default value of an uninitialized std::array variable is
built element by element from N, a struct with a std::array field does not
derive Default, and it does not derive Copy either since the field is a
Vec.
Warning
An empty initializer,
std::array<int, 3> a = {};, becomesvec![], a vector of length 0, where C++ value-initializesNelements; indexing it panics (#313).
std::string and streams
std::string maps to Vec<libc::c_char> in the unsafe model and Vec<u8> in
the refcount model through its rules; the converter itself only special-cases
string literals (their type in an initializer, and ASCII escaping) and
range-for over a string. std::ostream calls (std::cout << x) are detected
with IsCallToOstream and translated by a dedicated path rather than by rules;
that path is described with printf under Expressions.