rhdl_ed25519_core/lib.rs
1//! # Dalek-compatible Ed25519 hardware in RHDL
2//!
3//! This crate is the hardware facade for the workspace. Its modules re-export
4//! the real SHA-512, field, scalar, point, message-feeder, and wire-type crates
5//! used by the synthesizable design. It contains no Dalek implementation code:
6//! the `rhdl_ed25519_model` crate uses the pinned Dalek snapshot only as a
7//! differential-test oracle.
8//!
9//! The workspace has two hardware tracks:
10//!
11//! - The **compatibility core** implements public-key derivation, cold and
12//! cached-key signing, strict verification, multipart byte streams, and key
13//! clearing. `rhdl_ed25519_top::Ed25519Core` assembles the modules exported
14//! here into the end-to-end synchronous design.
15//! - The **fixed-64 fast path** is an optimization track for hardware key
16//! expansion followed by saturated cached-key signing of 64-byte messages on
17//! an AMD Alveo U280. Its `fast_field`,
18//! `fast_fixed_base`, `fast_sha512`, `fast_point_codec`, `fast_scalar`, and
19//! `fast_sign` crates contain the specialized SystemVerilog datapath, RHDL
20//! wrapper types, cycle models, and RTL tests. It is not yet a replacement for
21//! the compatibility core.
22//!
23//! **Implementation boundary:** the optimized fixed-64 datapath is handwritten
24//! SystemVerilog and is compiled directly by Icarus, Verilator, and Vivado. It
25//! is not emitted from RHDL. The
26//! [SystemVerilog backend manual](../ed25519_fast_sv_docs/index.html) documents
27//! every module, parameter, helper function, port group, state machine, memory
28//! image, and build boundary in that rewrite.
29//!
30//! ## End-to-end signing
31//!
32//! For an RFC 8032 seed `sk` and message `M`, the hardware executes:
33//!
34//! ```text
35//! SHA512(sk) -----------------> clamp -> a -------> [a]B -> encode -> A
36//! |
37//! +-----------------------> prefix
38//!
39//! SHA512(prefix || M) --------> reduce mod l -> r -> [r]B -> encode -> R
40//! SHA512(R || A || M) --------> reduce mod l -> k
41//! S = r + k*a mod l
42//! signature = R || S
43//! ```
44//!
45//! The host never supplies a prehash, expanded key, reduced nonce, challenge,
46//! or intermediate point. SHA-512 padding, clamping, reduction, scalar
47//! multiplication, point compression, and signature assembly all execute in
48//! hardware.
49//!
50//! ## Fixed-64 pipeline anatomy
51//!
52//! The cached fixed-64 signer is a tagged task graph, not a 512-stage linear
53//! pipeline. `512` is the number of requests in the standard throughput test.
54//! The physical concurrency is 64 live signing contexts, 16 interleaved point
55//! contexts, a three-phase SHA-512 round datapath reused for 80 rounds, and a
56//! 17-position data/valid pipeline in each radix-51 field multiplier. A SHA worker
57//! occupies one block for 242 cycles even though it has only three round
58//! phases; worker occupancy and physical register depth are different values.
59//!
60//! One cached signing request advances through these dependency stages:
61//!
62//! | Stage | Function | Main hardware and completion state |
63//! |---|---|---|
64//! | 1. Allocate | Accept a 64-byte message and caller tag into a free context | An eight-context rotating search writes banked context registers and raises `nonce_hash_pending` |
65//! | 2. Build nonce block | Form `prefix || M` with SHA-512 padding and length | One of two builders emits 16 words and submits a context-tagged nonce request |
66//! | 3. Hash nonce | Execute all 80 SHA-512 rounds | One of four three-phase compression workers returns a 512-bit digest through a nonce FIFO |
67//! | 4. Reduce nonce | Compute canonical `r = digest mod l` | The dedicated 626-DSP, 46-position nonce reducer writes `nonce_scalar` and raises `point_r_pending` |
68//! | 5. Compute nonce point | Compute `[r]B` | The 16-context signed multi-comb engine performs 32 constant scans/mixed additions and seven doublings, then enqueues projective `(X,Y,Z)` |
69//! | 6. Compress nonce point | Encode projective `[r]B` as Edwards-Y `R` | One of four pair-codec ways shares one inversion across two points, stores `R`, and raises `challenge0_pending` |
70//! | 7. Hash challenge block 0 | Compress the full 128-byte `R || A || M` block | A SHA worker saves the returned chaining state in the requesting context and raises `challenge1_pending` |
71//! | 8. Hash challenge block 1 | Compress the padding-only block for a 1024-bit input | A SHA worker emits the final challenge digest through a challenge FIFO |
72//! | 9. Reduce challenge | Compute canonical `k = digest mod l` | The second 626-DSP, 46-position reducer writes `challenge_scalar` and raises `muladd_pending` |
73//! | 10. Compute `S` | Compute canonical `S = r + k*a mod l` | A 69-position, 1,237-DSP multiply-add pipeline returns `S` with the context tag |
74//! | 11. Retire | Assemble `signature = R || S` and restore the caller tag | The output pulses valid and releases the signing context |
75//!
76//! Pending bits and bounded FIFOs form the scoreboard between stages. Workers
77//! can therefore service different requests simultaneously and complete
78//! internal stages out of input order; the six-bit context tag associates every
79//! completion with its retained state. The AXI wrapper adds a 64-entry result
80//! FIFO because the core result is a one-cycle pulse. See [`architecture`] for
81//! the exact eligibility rules, queue topology, key-load path, backpressure,
82//! banking, and stage-by-stage source map.
83//!
84//! ## Strict verification
85//!
86//! Verification rejects malformed inputs before evaluating the group equation:
87//!
88//! ```text
89//! canonical(S)?
90//! decode(A) -> reject malformed or small-order A
91//! decode(R) -> reject malformed or small-order R
92//! SHA512(R || A || M) -> reduce mod l -> k
93//! compare R with encode([S]B - [k]A)
94//! ```
95//!
96//! Dedicated result codes distinguish framing errors, a missing cached key,
97//! malformed points, non-canonical `S`, small-order points, and equation
98//! failure. See [`types`] for the wire-level contract.
99//!
100//! ## Arithmetic hierarchy
101//!
102//! The compatibility backend follows Dalek's serial 32-bit organization:
103//!
104//! - [`field`] represents values modulo `p = 2^255 - 19` with ten alternating
105//! 26/25-bit limbs stored in 32-bit lanes.
106//! - [`scalar`] represents values modulo the Ed25519 group order `l` and
107//! implements wide reduction, 256-bit reduction, canonicality, and
108//! `r + k*a mod l`.
109//! - [`point`] uses extended Edwards coordinates `(X:Y:Z:T)`, Projective Niels
110//! points, balanced radix-16 recoding, and fixed-pattern candidate scans.
111//! - [`sha512`] implements padding, a 16-word circular schedule, and all 80
112//! compression rounds in synchronous logic.
113//! - [`hash_feeder`] inserts hash prefixes and retains the first 4 KiB of a
114//! message for the second signing pass.
115//!
116//! ## Fast-path component and primitive map
117//!
118//! The current matching-source U280 optimized out-of-context core report
119//! divides into seven direct children. Registers, RAMB36s, and DSPs sum exactly.
120//! Vivado combines 480 LUTs across hierarchy boundaries, so the adjusted LUT
121//! total is smaller than the raw direct-row sum. A nested field multiplier must
122//! not be added to its parent a second time.
123//!
124//! | Component | Function | LUTs | Registers | RAMB36 | DSP48E2 |
125//! |---|---|---:|---:|---:|---:|
126//! | Top-level context/control | Own 64 contexts, banked state, scoreboards, arbiters, tags, and queues | 10,126 | 8,096 | 0 | 0 |
127//! | Four-worker SHA pool | Build and compress seed, nonce, and challenge blocks | 18,489 | 12,088 | 4 | 48 |
128//! | Nonce reducer | Reduce the 512-bit nonce digest modulo `l` | 2,722 | 11,982 | 0 | 626 |
129//! | Challenge reducer | Reduce the 512-bit challenge digest modulo `l` independently | 2,844 | 11,987 | 0 | 626 |
130//! | Fixed-base point engine | Compute `[a]B` and `[r]B` with signed multi-comb arithmetic | 82,095 | 84,605 | 352 | 2,116 |
131//! | Four-way point codec | Batch-invert and compress projective nonce points | 84,295 | 67,702 | 0 | 2,092 |
132//! | Scalar multiply-add | Compute and canonicalize `r + k*a mod l` | 2,253 | 23,343 | 0 | 1,237 |
133//! | **Raw direct-row sum** | | **202,824** | **219,803** | **356** | **6,745** |
134//! | **Adjusted core total** | | **202,344** | **219,803** | **356** | **6,745** |
135//!
136//! SHA uses LUTs for Boolean/rotate logic, FFs for working words and schedules,
137//! one RAMB36 constant ROM per worker, and 12 DSPs per worker for marked 64-bit
138//! state/add operations. Each 512-bit scalar reducer uses 626 DSPs. Every
139//! radix-51 field multiplier uses 523 DSPs and has a 17-position,
140//! initiation-interval-one pipeline: four in the point engine and four in the
141//! codec account for 4,184 DSPs. Two field add/sub lanes add 24 DSPs. Scalar
142//! arithmetic accounts for 2,489 DSPs: two 626-DSP digest reducers and the
143//! 1,237-DSP multiply-add hierarchy. The remaining 48 DSPs are in SHA. The 356
144//! RAMB36 total is exactly 352 multi-comb table memories plus four SHA constant
145//! ROMs. [`architecture`] records LUTRAM/SRL use, timing paths, memory traffic,
146//! placement boundaries, and evidence filenames.
147//!
148//! ## Memory and traffic
149//!
150//! Arithmetic engines use registers and local BRAM/ROM only. They do not use
151//! HBM or DDR as scratch storage. For a signing message of `N` bytes, the 4 KiB
152//! cache gives the logical external message traffic
153//!
154//! ```text
155//! N + max(N - 4096, 0)
156//! ```
157//!
158//! before AXI-line padding. Verification reads `N` message bytes once. The
159//! top-level result separately reports logical stream bytes and external read
160//! bytes; host/AXI/HBM traffic must also count command records, result records,
161//! line padding, and data-mover behavior.
162//!
163//! ## Secret-dependent work
164//!
165//! Scalar loops have fixed trip counts. Point-table lookup reads every candidate
166//! bank at the same public address and selects only after the reads. A secret
167//! digit must not select a BRAM address, enabled bank, iteration count, or stall
168//! pattern. Zero digits execute the same point-operation schedule as nonzero
169//! digits. These properties require emitted-RTL address and enable tests in
170//! addition to functional Rust tests.
171//!
172//! ## Source map
173//!
174//! The modules below are the stable way to navigate arithmetic and protocol
175//! APIs. Their items are re-exported from the owning implementation crates.
176//! Higher-level control is intentionally split into separate crates:
177//!
178//! | Crate | Responsibility |
179//! |---|---|
180//! | `rhdl_ed25519_top` | child wiring, arbitration, and top-level synchronous design |
181//! | `rhdl_ed25519_controller_types` | top-level I/O, states, and status |
182//! | `rhdl_ed25519_transition` | pure high-level next-state kernel |
183//! | `rhdl_ed25519_commands` | child-engine command generation |
184//! | `rhdl_ed25519_registers` | retained cache/work/point register updates |
185//! | `rhdl_ed25519_result` | result and error classification |
186//! | `rhdl_ed25519_api` | Dalek signing traits over a driver/simulator transport |
187//! | `rhdl_ed25519_model` | pinned-Dalek reference behavior and vectors |
188//! | `rhdl_ed25519_sim` | cycle simulation, traces, RTL export, and checksums |
189//! | [`architecture`] | current backend map, pipeline, cycle breakdown, and contributor guide |
190//!
191//! ## Evidence and performance status
192//!
193//! The compatibility core has end-to-end RHDL simulation evidence, including
194//! exact RFC 8032 signing and strict verification. The recorded 64-byte signing
195//! run takes 90,928 core cycles. A U280 out-of-context synthesis constrained to
196//! 300 MHz records 237,294 LUTs, 94,129 registers, 800 DSPs, and 6 block-RAM
197//! tiles, but has negative WNS and therefore must not be described as a closed
198//! 300 MHz implementation.
199//!
200//! The current cached fixed-64 source passed a 512-signature RTL benchmark
201//! against a Dalek-derived public key and signature checksum. The first-to-last
202//! output span is 101,438 cycles across 511 intervals, or 198.508806
203//! cycles/signature. One hardware key load plus fill/drain gives a 119,459-cycle
204//! finite batch, or 233.318359 cycles/signature.
205//!
206//! Matching-source U280 optimized OOC core synthesis uses 202,344 LUTs,
207//! 219,803 registers, 356 RAMB36 tiles, 6,745 DSPs, and no URAM. It is 57,656
208//! LUTs, 300,197 registers, and 2,255 DSPs below the project budgets. At the
209//! 200 MHz request, WNS is +0.520 ns. The 4.462 ns worst path is inside the
210//! scalar reducer, so this report does not establish a 4.000 ns/250 MHz clock.
211//! Combining 200 MHz with the RTL interval projects 1,007,512 sustained cached
212//! signatures/s and 857,198 signatures/s for the finite batch. These are
213//! simulation-plus-OOC projections, not routed or measured FPGA throughput.
214//!
215//! Physical implementation currently belongs to the preceding 3,376-DSP
216//! builder-pipeline source. Its standalone core is fully routed and closes
217//! setup at 199 MHz with +0.018 ns WNS. Its full U280 platform is also fully
218//! placed and routed, but misses setup by 1.836 ns with 137,269 failing
219//! endpoints; Vitis therefore emits no bitstream or xclbin. Seven active RTL
220//! files and the host source differ from that predecessor package. The current
221//! 6,745-DSP source has not been placed or routed, so predecessor timing must
222//! not be combined with current cycles as a routed claim.
223//!
224//! The completed U280 result belongs to the preceding single-scan digest-FIFO
225//! revision. Its packaged 220 MHz xclbin closed with +0.001 ns WNS and measured
226//! 201,933 steady signatures/s (p50) over a 4,160-signature batch, five warm-ups,
227//! and 30 runs; output checksums matched the Dalek-derived reference. The
228//! packaged kernel synthesis used 210,098 LUTs, 101,849 registers, 90 RAMB36,
229//! and 720 DSPs. It is historical evidence and must not be attributed to the
230//! current cached pipeline. See [`architecture`] for the current hierarchy,
231//! evidence matrix, memory analysis, and AMD comparison status. A matching
232//! fixed-64 AMD wrapper now has the same AXI ABI and HBM bank map, but neither
233//! implementation has a current comparable U280 hardware benchmark.
234//!
235//! ## Contributor workflow
236//!
237//! The repository-level `CONTRIBUTING.md` is the complete workflow. In short:
238//!
239//! 1. Start in the owning module below, not in generated Verilog.
240//! 2. Add an independent arithmetic or Dalek differential test first.
241//! 3. Compile the affected RHDL kernel and simulate emitted RTL when lowering or
242//! a SystemVerilog black box is involved.
243//! 4. Inspect secret-dependent address, enable, and valid traces.
244//! 5. Run the end-to-end simulator before changing a compatibility claim.
245//! 6. Regenerate Vivado evidence before changing an area or timing claim.
246//!
247//! On memory-constrained hosts, use one Cargo job and iterate on narrow crates:
248//!
249//! ```text
250//! cargo test -p rhdl_ed25519_model --release -j 1
251//! cargo test -p rhdl_ed25519_core --release -j 1
252//! cargo test --workspace -j 1
253//! cargo doc --workspace --no-deps -j 1
254//! ```
255
256#![allow(clippy::needless_range_loop)]
257
258#[doc = include_str!("../../../docs/ARCHITECTURE.md")]
259pub mod architecture {}
260
261pub mod field {
262 //! Field arithmetic modulo `2^255 - 19`.
263 //!
264 //! The compatibility backend uses [`FieldElement2625`] and exposes serial
265 //! and four-lane engines. Encoding, inversion, square-root chains, and
266 //! parallel multiplication are implemented by the owning
267 //! `rhdl_ed25519_field` crate.
268
269 pub use rhdl_ed25519_field::*;
270}
271
272pub mod hash_feeder {
273 //! Streaming SHA-512 input assembly and the 4 KiB message replay cache.
274 //!
275 //! [`HashFeederEngine`] combines an internal [`crate::sha512::Sha512Engine`]
276 //! with prefix selection, 64-bit message-word framing, logical/external byte
277 //! counters, and signing-pass replay requests.
278
279 pub use rhdl_ed25519_hash_feeder::*;
280}
281
282pub mod point {
283 //! Edwards-point arithmetic, scalar multiplication, and point encoding.
284 //!
285 //! This module re-exports the point controller, balanced-radix-16 scalar
286 //! multiplier, strict compression/decompression controller, and the
287 //! constant-time 32-byte equality kernel used by verification.
288
289 pub use rhdl_ed25519_bytes_equal::*;
290 pub use rhdl_ed25519_point::*;
291 pub use rhdl_ed25519_point_codec::*;
292 pub use rhdl_ed25519_scalar_mul::*;
293}
294
295pub mod scalar {
296 //! Scalar arithmetic modulo the Ed25519 group order.
297 //!
298 //! [`ScalarEngine`] accepts wide and 256-bit reductions, modular multiply-add,
299 //! and canonicality commands through one fixed-latency state machine.
300
301 pub use rhdl_ed25519_scalar::*;
302}
303
304pub mod sha512 {
305 //! End-to-end hardware SHA-512.
306 //!
307 //! [`Sha512Engine`] accepts an optional internal prefix and a byte stream,
308 //! constructs FIPS 180-4 padding in hardware, and emits the eight-word
309 //! digest. No host prehash is part of the public path.
310
311 pub use rhdl_ed25519_sha512::*;
312}
313
314pub mod types {
315 //! Public command, message-stream, pass-request, result, flag, and error
316 //! encodings shared by the core, simulator, and FPGA shell.
317
318 pub use rhdl_ed25519_types::*;
319}
320
321pub use types::*;