Fidelity ledger
For each model sparx implements: the references it is checked against, every difference found between a reference and the science or another reference, what sparx does, and the test that pins it. design.md describes the tiers. A row is added before a model ships, not after.
Reference versions: SpikingJelly at commit c6cb8e46 (2026-10-03), snnTorch 1.0.0, DCLS 0.1.1 (Hammouamri et al.’s layer), IGITUGraz/eligibility_propagation at its default branch (Bellec et al.), Thvnvtos/SNN-delays at commit d169b4e3 on the 2023 SpikingJelly tree it was written against (Hammouamri et al.), all with torch 2.14.1+cpu; NEST 3.10.0 and Brian2 2.10.1 for the biophysical models. The fixture tools under tools/ record the versions in each fixture.
Neurons
Section titled “Neurons”| Model | Reference | Checked | Departures and choices |
|---|---|---|---|
LIF (sparx.dynamics.LIFCell, nn.LIF) | SpikingJelly LIFNode(decay_input=False); snnTorch Leaky(reset_delay=False) | Spikes exact and gradients within 3.6e-7 of both, soft, hard and detached resets (tests/test_reference.py) | Sparx’s decay is exp(-1/tau), the exact decay of the leak over a step; SpikingJelly’s is 1 - 1/tau, the Euler step. The model takes the decay itself, so either convention is a value, and the parity tests pass SpikingJelly’s. |
| LIF reset timing | snnTorch’s default reset_delay=True | Pinned as different (test_snntorchs_default_delayed_reset_is_a_different_model) | snnTorch’s default subtracts the threshold one step after the spike, without decaying it. That is no discretization of the continuous LIF, whose reset happens at the spike and then decays. Sparx resets on the spiking step. |
| LIF after an overshoot | snnTorch reset_delay=False | Pinned (test_snntorch_loses_spikes_after_an_overshoot_where_sparx_fires) | When a soft reset leaves the membrane at or above threshold, snnTorch subtracts the next reset before testing for a spike, so the neuron needs twice the threshold to fire again and loses spikes. Every neuron where the two diverge does so on the step after its own overshoot. Sparx and SpikingJelly fire whenever the membrane is at threshold, removing one threshold per spike. |
| LIF spike condition | snnTorch fires on v - threshold > 0; SpikingJelly and sparx on >= 0 | Fixtures keep every membrane 1e-4 from threshold | Differs only at exact equality. |
| Detached reset | snnTorch detaches its delayed reset and keeps the gradient of its immediate one | Gradients match with detach_reset=False | Sparx’s detach_reset is SpikingJelly’s option of the same name. |
Current-based LIF (Serial(LICell, LIFCell), nn.Synaptic) | snnTorch Synaptic(reset_delay=False) | Spikes exact and gradients matched, soft and hard reset | Same reset timing and overshoot departures as LIF. The synaptic current is an LICell whose membrane feeds the LIFCell each step, the same operations in the same order as snnTorch’s single cell. |
Adaptive LIF (sparx.dynamics.ALIFCell) | Bellec et al. 2020 equations; their code, Figure_4_and_5_ATARI/alif_eligibility_propagation.py | Spike for spike against a transcription of their cell with a reset of decay * threshold (test_alif_is_bellecs_model_with_its_reset_decayed) | A spike subtracts the baseline threshold, as theirs does (sparx subtracted the adaptive threshold before this audit). Their reset lands one step later, undecayed, so their model with reset decay * threshold is sparx’s. Their repository has a second variant, LightALIF, that scales input by 1 - decay and adaptation by 1 - adapt_decay; sparx follows the paper’s equations, which the ATARI cell implements. Refractoriness is their counter (refractory, their n_refractory, a duration of round(refractory / dt) steps): spike for spike against a transcription of it (test_alif_refractoriness_is_bellecs_counter), and off by default. Their pseudo-derivative is Triangle(width=threshold, scale=0.3 / threshold) on v - A; sparx’s default surrogate is ATan for every neuron. |
Escape-noise LIF (sparx.dynamics.BernoulliCell) | Its equations, the discrete-time escape noise of Pfister et al. 2006: the LIF membrane, firing with probability sigmoid(beta (v - threshold)) per step | Spike for spike against a float64 loop given the same uniform draws, each reset rule, probabilities within 1.9e-7 (test_bernoulli_matches_the_reference_loop); noise 1 - s replays the spikes s; a run in two chunks equals one | The noise is an input (SynapticInput.noise), drawn by the caller, so the cell is a pure function and a trajectory can be forced. p is per step whatever dt is. The spike is a sample and passes no gradient. |
RNeuralNet’s neuron (sparx.dynamics.PulseCell) | Soma_t::ActivationFunction of RNeuralNet-Research (AshishKumar4/RNeuralNet-Research d4b7803), A_CONST 1 | Against the original program (RNeuralNet under learning rules) and its formula in float64 (test_a_pulse_cell_is_rneuralnets_activation, observed 7.0e-8) | The original’s sum empties once the neuron has sent on all its connections; in the order RNeuralNet runs it, that is every tick, so the cell keeps only a jump that lands after its step. exp(min(s - threshold, 0)) - 1 keeps the unused branch finite and equals the original’s below the threshold. |
Izhikevich (sparx.dynamics.Izhikevich, scheme="published", order="izhikevich", the defaults; nn.Dynamics(Izhikevich())) | Izhikevich 2003, the paper’s MATLAB code | Spike for spike in float64 against a transcription of the code (test_izhikevich_matches_his_published_loop_spike_for_spike_in_float64); regular spiking under input 10 adapts, then fires every 47 to 62 ms | Each 1 ms step takes two half-steps of v and one step of u, as the code does (sparx took 0.5 ms Euler steps of both before this audit). The firing patterns the model is known for are those of this scheme; an integration closer to the ODE is a separate, named model (design.md section 4.2). The peak is not clipped at 30 mV, as in the 2003 code. |
| Izhikevich numerics | XLA | Measured | Compiled XLA rounds fused arithmetic differently from NumPy in the last bit (22% of elements of one step). The quadratic membrane amplifies that chaotically, so exact comparisons run op by op (jax.disable_jit). |
| PSN, MaskedPSN, SlidingPSN | SpikingJelly neuron/psn.py | Spikes exact, gradients within 1.2e-6 | None found. MaskedPSN’s lambda_ is the call argument masking. |
Biophysical LIF (sparx.dynamics.LeakyIntegrateAndFire) | Its equation’s exact solution over a step; NEST iaf_psc_exp, iaf_psc_alpha, iaf_psc_delta; Brian2 exact | Ground truth in float64 (tests/test_dynamics.py): one 5 ms step equals 500 steps of 0.01 ms to 1e-12, conductance relaxation and f-I periods are the analytic ones. Against NEST (tests/test_simulators.py): spike for spike and voltages within 1e-11 mV over 300 ms of random weighted input, for exponential, alpha and delta synapses | None for current input. Refractoriness holds round(t_ref / dt) steps after the spiking step, as NEST does. |
| Refractoriness | Brian2 | Spike for spike and within 1e-9 mV once sparx’s t_ref is one step shorter (test_brian2_*) | Brian2 counts t_ref from the start of the step in which the neuron crossed, the time it stamps on the spike, so it holds one step less than NEST. The crossing happened within that step, so Brian2’s period is up to a step shorter than t_ref and NEST’s up to a step longer; sparx follows NEST, which never undercuts the stated period. |
| Conductance-based LIF | NEST iaf_cond_exp, iaf_cond_alpha, iaf_cond_beta (adaptive RK45); Brian2 exponential_euler; a float64 RK4 at dt / 10 | NEST is within 1.5e-9 mV of the RK4 truth. Sparx’s default (hold="mean") is within 2e-3 mV of NEST and fires with it spike for spike until a crossing within that margin; its error falls fourfold when dt halves (test_conductance_hold_converges_at_second_order_to_the_rk4_truth). hold="start" is Brian2’s scheme: spike for spike and within 1e-9 mV | Within a step the membrane sees a held conductance. Brian2 holds its value at the start of the step, first order, and fires several steps late against NEST over 300 ms. Sparx holds its exact average over the step, which integrates the leak’s decay exactly and is second order, at the same cost. NEST’s adaptive integration is more accurate still and does not vectorize over a population. |
AdEx (sparx.dynamics.AdEx) | Brette and Gerstner 2005; NEST aeif_psc_exp, aeif_cond_exp (adaptive RK45); Naud et al. 2008’s firing patterns | Each of Naud’s eight parameter sets for 500 ms: the same spike count as NEST and every spike within 0.8 ms (ten substeps), within 0.1 ms (a hundred); the irregular set is chaotic, and its spike count and irregularity match. The published patterns hold (tonic, adapting, initial burst, regular bursting, delayed accelerating, delayed regular bursting, transient). With synaptic input, spike for spike within 0.2 ms (current) and 0.5 ms (conductance) | NEST’s right-hand side reads the voltage capped at v_peak, and v_reset while refractory; sparx reads it the same way. NEST resets inside its adaptive steps and integrates the rest of the step from the reset, and sparx does so at the substep that crosses: without that, spikes drift 5 ms late over 500 ms whatever the substep count. The upswing is stiff, so the default is ten RK4 substeps per step; an exponential Rosenbrock step was less accurate than RK4 at every count tried. |
Izhikevich in physical time (sparx.dynamics.Izhikevich) | Izhikevich 2003; NEST izhikevich, both schemes | With order="nest", op by op in float64, NEST’s run to the last bit for 300 ms at 1 ms (test_izhikevich_published_scheme_is_nests_to_the_last_bit); compiled at 0.1 ms, each of the paper’s seven classes fires as many spikes as NEST, every one within 1 ms, under both schemes; delta input spike for spike within 1e-9 mV | NEST starts u at -13 whatever b is; sparx starts it at b * v, as Izhikevich’s code does, and the fixtures set NEST’s U_m so. NEST computes (0.04 v) v, Izhikevich’s code 0.04 (v^2); the two differ in the last bit, and the field order names which one a model computes: "izhikevich" by default, "nest" in the NEST tests. NEST adds a delta input to the voltage under Euler and as a one-step current of the weight under the published scheme; sparx’s jump is always a voltage jump, and the published convention is the weight passed as current. |
Izhikevich (2004)‘s twenty firing patterns (izhikevich_2004, scheme="semi_implicit") | His figure1.m (archived copy), run unchanged but for its plotting in GNU Octave 8.4 (tools/make_izhikevich_2004_fixtures.py) | Each of the twenty panels with his current, step and start: every spike on his step, and the voltage between spikes his to a relative 1e-9 in float64, compiled (test_izhikevich_2004_patterns_are_his_codes_runs) | His code takes one Euler step of v and then one of u from the new v, which is neither the 2003 code’s scheme nor NEST’s, so it is a third, named scheme. Two panels change the equations: class 1 excitability and the integrator use 4.1 v + 108, accommodation du/dt = a b (v + 65); these are fields (quadratic, v_u, u_decay), not hidden cases. In the class 2 panel the voltage differs by up to 0.012 mV and no spike moves: Octave’s V^2 is libm’s pow, which rounds differently from v * v in about one value in a thousand, and that panel’s slow ramp through the bifurcation amplifies it. |
Hodgkin-Huxley (sparx.dynamics.HodgkinHuxley) | Hodgkin and Huxley 1952 in NEST hh_psc_alpha’s statement (rates, 100 pF, peak detection, 2 ms refractoriness); Brian2 exponential_euler | Four currents from rheobase to 4 nA for 300 ms: every spike within one step of NEST’s (default), spike for spike and within 0.02 mV with rk4 at 0.025 ms; with alpha synapses under inhibition down to -128 mV, spike for spike. exponential_euler is Brian2’s within 1e-10 mV for 10 ms. The default scheme converges at second order (test_hodgkin_huxley_strang_is_second_order) | NEST integrates by adaptive RK45. Fixed-step RK4 is more accurate per cost but diverges once inhibition makes the gates stiff (beta_m reaches 130/ms at -128 mV), and in a network that is not rare; sparx’s default is a Strang splitting of exact gate and voltage solves, stable at any step. Brian2’s exponential_euler (Rush-Larsen gates, exact voltage, first order) fires up to 1.6 ms off NEST at 0.1 ms. A spike is the voltage’s peak, reported at the step where it has begun to fall, as NEST does. |
NMDA magnesium block (MgBlock) | Jahr and Stevens 1990 | The formula (test_mg_block_is_jahr_and_stevens); applied at the voltage at the start of the step | Held over the step like the conductance it scales. |
Synapses and layers
Section titled “Synapses and layers”| Layer | Reference | Checked | Departures and choices |
|---|---|---|---|
Exponential, Alpha, BiExponential, Delta synapses | NEST’s iaf_psc_* and iaf_cond_* kinetics; peaks normalized to the weight, as NEST does | Through the neuron tests above, and analytically: decay, peak value and time (test_dynamics.py). The response of the membrane to each waveform is the exact integral, checked against adaptive quadrature to 1e-12 at equal, close and distant time constants | Spikes land at the end of their arrival step and shape the membrane from the next, NEST’s order (an event stamped T lands at the end of the step ending at T). Brian2’s on_pre with no delay lands at the same point. A synapse declares where its arrivals land (lands): Delta before the threshold test, as NEST’s iaf_psc_delta input, or after it with after_threshold=True, as Brian2’s on_pre="v += w", through the neuron model’s after_threshold, which drops the jump where a reset sets the membrane and keeps it under a subtracting reset or none. |
DelayedDense | DCLS Dcls1d (version gauss in training, max with rounded positions at evaluation), as SNN-delays uses it | Outputs within 2.4e-7, gradients (input, weight, delay) within 1.4e-6, both modes (test_delayed_dense_matches_dcls_delays) | Sparx’s sigma is DCLS’s effective width, abs(SIG) + 0.27, and its delay is K - 1 - (P + K // 2). SNN-delays also pads the input on the right by (K - 1) // 2, which lengthens the output so delayed spikes after the input’s end reach the readout; the layer keeps the length (it is causal and streams exactly), and SpikingMLP(extend=True) appends those zeros to each delayed synapse’s input, their setting. DCLS clamps its positions after every update and has no clip in its kernel, so a delay at an end of the range gets the kernel’s full gradient; sparx’s kernel clips the center with the gradient passed straight through (a plain jnp.clip halves it at the end, which the network test below caught), and dew’s ParamGroup(bounds=...) clamps after each update. Their width decays exponentially to SIG = 0.23 (effective 0.5) over the first quarter of training, dew’s Exponential(init=12, end=0.23, decay_steps=epochs // 4, offset=0.27) for their 25-step kernel; the alif example uses a linear schedule to 0.5. |
Delays per connection (sparx.dynamics.Sparse with delay and longest_delay, in a RecurrentCell) | Messages queued on each connection and weighted when sent, as RNeuralNet-Research’s Neurite_t queues them | Against a float64 loop that queues each message for its delay (test_a_delayed_sparse_recurrence_is_the_queued_loop, observed 4.8e-7 on outputs up to 7.4; without the delays the outputs differ by more than 0.1), and through RNeuralNet against the original program; a stream fed in chunks equals one run; fast weights on delayed connections against a float64 loop (test_fast_weights_on_delayed_connections_pair_what_each_one_delivers) | A message keeps the weight it left with, so a change reaches only later messages. On a delayed wiring, fast weights pair each connection’s new output with what it delivers this step, its source’s output delay steps before (RecurrentState.history), and a message carries the fast weights of the step that sent it. |
| SEW ResNet blocks | SpikingJelly model/sew_resnet.py BasicBlock | Spike for spike in float64 for ADD, AND and IAND, with and without the downsampling shortcut (test_sew_block_matches_spikingjelly) | None found. The small stem (one 3x3 convolution, no pooling) is sparx’s addition for 32x32 inputs; the imagenet stem is SpikingJelly’s. |
TsodyksMarkram | Tsodyks and Markram 1997; NEST tsodyks2_synapse | The weight transmitted at every presynaptic spike, depressing and facilitating parameter sets, four synapses for 2 s: within 1e-10 relative (tests/test_plasticity.py) | NEST transmits its initial w u x at the first spike and its initial u is 0.5 whatever U is; a synapse at rest has u = U and transmits w U, which sparx does, and the fixtures set NEST’s u so. |
PairSTDP | Guetig et al. 2003 (additive at mu = 0, Song et al. 2000; multiplicative at mu = 1); NEST stdp_synapse | Every transmitted weight within 1e-10 relative, additive and multiplicative, with a 1 ms dendritic delay | None. NEST evaluates the rule at presynaptic spikes from the postsynaptic history; sparx applies the same updates as the spikes happen, in the same order, including NEST’s for coincident spikes (a spike pairs only with strictly earlier spikes of the other side, and potentiation precedes depression). |
TripletSTDP | Pfister and Gerstner 2006, all-to-all; NEST stdp_triplet_synapse | Every transmitted weight within 1e-10 relative | None. The defaults are the paper’s visual cortex fit; NEST keeps the postsynaptic time constants on the neuron, with defaults of 20 and 110 ms rather than the fit’s 33.7 and 125 ms. |
DopamineSTDP | Izhikevich 2007; NEST 3.10 stdp_dopamine_synapse (Potjans et al. 2010) with a volume_transmitter | Every transmitted weight within 1e-10 relative (observed 2.1e-12 to 4.7e-12) for four synapses over 2 s with a dopamine neuron arriving 1 ms after it fires: NEST’s defaults, which sink to Wmin; a baseline b = 0.01 under which n - b changes sign between dopamine spikes, unclipped; and bounds reached on both sides (test_dopamine_stdp_is_nests). In a Network, the weights learned equal the rule run alone on the network’s own spikes and dopamine trace (test_dopamine_stdp_in_a_network_is_the_rule_on_the_networks_own_spikes). Replacing the step’s exact integral by c (n - b) dt fails the NEST test. | NEST integrates between events and clips at each; sparx integrates the same exact solution step by step and clips after every step. They differ only when n - b changes sign between dopamine releases and the weight crosses a bound in that interval. The rule’s tau_n is a Python number the network checks against its modulator’s tau. |
Networks
Section titled “Networks”| Network | Reference | Checked | Departures and choices |
|---|---|---|---|
Recurrent Network (sparx.graph) | NEST, the same edges, weights and per-edge delays (1 to 3 ms) | 60 neurons, 10% connectivity, mixed-sign weights, per-neuron currents, 500 ms: spike for spike with iaf_psc_exp and iaf_psc_delta networks; with iaf_cond_exp, spike for spike until a held-conductance difference of 1e-3 mV moves one spike by one step (after 50 ms), and total activity within 5% (tests/test_graph.py) | A spike sent in step m over D steps lands at the end of step m + D, NEST’s and Brian2’s timing for a delay of D dt. Kinetic synapses accept D = 0 (Brian2’s default); delta synapses need D >= 1. A delay, duration or chunk that is not a whole number of steps raises (sparx.graph.network.whole_steps). Rounding it would move it by up to half a step, in a direction set by floating-point error. |
Brunel (2000), sparx.graph.models.brunel | NEST’s brunel_delta_nest.py, the same parameters and connection rule | At 2,500 neurons, in each of Brunel’s four regimes, the excitatory rate, mean interspike CV and population Fano factor of one sparx run fall within NEST’s spread over eight seeds (test_brunel_regimes_match_nests_statistics) | At this size (in-degree 200) the regimes are not those of Brunel’s Figure 8, which needs his in-degree of 1,000; the parity holds either way, and NEST at full size gives the same. The slow synchronous regime varies by 25% between seeds in both simulators, so three seeds were not enough to compare it. |
Potjans and Diesmann (2014), sparx.graph.models.microcircuit | NEST 3.10 running the PyNEST implementation of INM-6/microcircuit-PD14-model (f79f8ac) unchanged, and running the network sparx draws, edge by edge (tools/make_microcircuit_fixtures.py) | At a fifth of the neurons and inputs (15,435 neurons, 12M synapses): the neurons, synapse counts and currents the reference derives, exactly, and weights and delays from its distributions; on the network sparx draws, NEST and sparx in float64 spike alike for 300 ms (12,689 spikes); over the reference’s own draws at 15 seeds, each population’s mean rate within 4 standard deviations of NEST’s, and the Kolmogorov-Smirnov distance of its rates, ISI CVs and pairwise correlations from NEST’s within the distances between NEST’s seeds, the reference’s own verification (Dasbach et al. 2021) (tests/test_microcircuit.py) | Weights and delays are drawn again where NEST’s redraw draws them again, as the reference’s code does; its model description clips them instead. At a tenth of the neurons and inputs the currents that make up for the lost inputs leave every population but L6I below threshold, and the network falls silent in NEST and sparx alike (L6I at 3.43 Hz in both), so the comparison runs at the reference’s own scale for its verification data. Over one second, layer 2/3’s excitatory rate varies between NEST’s seeds from 0.55 to 0.64 Hz. In float32 the runs part after about 140 ms, as rounding steers a chaotic network. |
CUBA and COBA, sparx.graph.models.cuba, coba | Brette et al. 2007’s benchmarks in Brian2: examples/CUBA.py, and its conductance-based counterpart | 4,000 neurons, 500 ms: both populations’ rates, the excitatory CV and Fano factor of one sparx run fall within Brian2’s spread over four seeds (test_brette_benchmarks_match_brian2s_statistics) | Brian2 counts refractoriness one step shorter (see Neurons). Sparx holds conductances at their step average, Brian2’s exponential_euler at the step start. Both act without delay, which sparx’s kinetic synapses take as Brian2 does. |
Shiu et al. (2024), sparx.graph.connectome.shiu2024 | Their Brian2 model (github.com/philshiu/Drosophila_brain_model), and their published runs on FlyWire v630 | Their neuron model on a small random graph: spike for spike with their Brian2 statement, through edge and event delivery (test_shius_neuron_model_fires_with_brian2_spike_for_spike). The whole brain (127,400 neurons, 14.7M connections), 21 sugar-sensing neurons activated, 10 trials against their 30: at 100 Hz, total spikes 9,675 vs 9,636, rates of the 347 neurons above 1 Hz correlated at 0.9989, MN9 67.1 Hz vs 67.0 +- 6.6; at 200 Hz, 16,909 vs 17,052, 0.9996 over 391 neurons, MN9 92.7 vs 93.3 +- 3.2 (test_shiu2024_reproduces_their_published_sugar_activation, run when their data is present) | Their model has quirks sparx reproduces as options, off elsewhere: a spike clears the synaptic variable g, and Brian2’s (unless refractory) on g both freezes it and drops input arriving while refractory; their stimulus lands after the threshold test (Delta(after_threshold=True)). freeze_synapses reads refractoriness from the neuron model (is_refractory), so it applies to any model; for one without refractoriness it drops only the arrivals of the step in which the neuron fired. Their published sugarR run stimulates at 200 Hz, as their notebook says, though their model.py now defaults to 150 Hz. |
Male CNS v0.9, Connectome.from_malecns | Janelia’s release tables | The reader, on synthetic tables (test_malecns_reader_keeps_neurons_and_signs_them_by_transmitter): 165,899 neurons (Traced and Anchor), 25.6M connections, 124M synapses from the v0.9 release | Shiu et al.’s model has no published run on this connectome, so it is checked against their FlyWire results. Their w_syn (0.275 mV) was fit to FlyWire v630’s synapse counts. The male CNS counts more synapses per neuron: a median of 347 input synapses against 206, and a mean of 748 against 414. At their weight the model runs away: the right sugar neurons (types LB3b, LB3c, LB3d and LB4b, the types of their 21) at 100 Hz recruit 18,249 neurons and 451,000 spikes a second. matched_w_syn scales the weight so the median neuron receives FlyWire’s input, which gives 0.163 mV. At that weight the same stimulus recruits 670 neurons and 11,500 spikes a second, against FlyWire’s 400 and 9,600, and drives MN9 (body 10331) at 81 Hz, against FlyWire’s 67 (test_shiu2024_on_the_male_cns_activates_mn9_from_sugar_neurons_at_the_matched_weight, which runs when the release tables are present). This is a calibration, not a reproduction: one scalar matched on one statistic, checked against one behaviour. The left sugar neurons drive MN9 at only 8 Hz, so the two sides differ, and FlyWire’s side labels are not checked against the male CNS’s here. |
Simulation over devices (simulate(mesh=MeshSpec(...))) | One device | On 4 simulated CPU devices of dew’s mesh: NEST’s 60-neuron network with its neurons partitioned over the data axis (15 per device) fires spike for spike as on one device over 300 ms; 8 trials spread over the data axis equal 8 trials on one; with MeshSpec(fsdp=2) the trials split over data and each trial’s neurons over fsdp, with the same spikes (tests/test_graph_distributed.py) | The step’s exchange of spikes between devices is left to the compiler’s partitioning of the gathers; a hand-written exchange of spike bitmasks (design.md 6.3) waits for measurements on real multi-device hardware. A population whose size the axis does not divide stays whole on every device, as dew’s logical_spec places any array. |
Resumed simulation (simulate(checkpoints=...)) | The unbroken run | A run stopped after 20 ms and resumed from its dew checkpoint to 50 ms gives the unbroken run’s spikes from 20 ms on and its final state, exactly (test_a_run_resumed_from_its_checkpoint_is_the_unbroken_run) | The checkpoint holds the state collection and the key, not the connectome or params, which the resumed run takes from the variables it is given. Records start at the checkpoint, and an accumulating monitor sums from there. |
Gradients through a Network | Central differences in float64; the same synapses delivered as an edge list | A network whose driven neurons fire, feeding neurons held below threshold through two trainable projections: the gradient of their membranes’ squared error equals central differences within 2.6e-9 relative (test_gradients_through_a_network_are_exact_where_no_membrane_nears_threshold). A trainable projection feeding a population with a recurrence delivered by events, with one delay or a delay per edge: the gradient equals the one through the same recurrence as an edge list within 8e-15 of the largest (test_gradients_through_event_delivery_are_the_edge_lists) | Event delivery is a while_loop, which has no reverse mode; it is differentiated as its edge list, so its backward pass visits every edge at every step. Stochastic release has no gradient. Where neurons fire the gradient is the surrogate’s, and detach_reset=True stops it through the reset. |
| Plastic and short-term projections | — | — | A plastic weight is read when a spike arrives; NEST reads it when the spike is sent, so a weight change while a spike is in flight (within the delay) is applied to it in sparx and not in NEST. Short-term release is computed when the spike is sent and travels with it, as in NEST. |
SNN-delays (Hammouamri et al., ICLR 2024), SpikingMLP with every synapse delayed | Their SnnDelays on DCLS 0.1.1 and the 2023 SpikingJelly their code runs on (tools/make_snn_delays_fixtures.py), a small copy of their SHD configuration: two hidden layers, delays of 0 to 4 steps on every synapse, the right padding, batch norm, decay_input=False LIF with hard reset and detach_reset, ATan(5), no bias, the non-spiking LIF readout and loss='sum' | A training step (Gaussian kernels, batch statistics): outputs and loss within 1.2e-7, every gradient (input, weights, delays, batch norm) within 5.7e-7, running means within 1e-5. An evaluation as their eval_model runs it (rounded delays, running statistics): outputs within 2.4e-7 (test_a_fully_delayed_network_trains_as_snn_delays, ..._evaluates_as_snn_delays) | torch’s running variance is the unbiased one and flax’s the biased one, a factor n / (n - 1) for n = steps x batch, which the test applies. Their LIF keeps 1 - 1 / tau of the membrane a step; sparx’s LIF keeps exp(-1 / tau'), so the recipe passes the tau' that gives their decay. Their dropout holds one mask over a recording (SpikingJelly’s multi-step Dropout), dropout_mask="sequence"; the fixture runs without dropout, since a random mask cannot be matched. |
Signalling beyond spikes
Section titled “Signalling beyond spikes”| Model | Reference | Checked | Departures and choices |
|---|---|---|---|
A model’s output (Output(value, offset), graded) | Every reference above | Every parity test in this ledger passes unchanged after spikes became Output.value; Shiu et al.’s whole-brain model fires 10,148 spikes in 1 s of sugar stimulation before and after, at the same 1.9 ms per step on CPU | LICell used to report its membrane as a spike; it is a graded model now, and sparx.nn.LI returns the same values. Event delivery reads spikes only and refuses a graded population. |
FLYNN (sparx.graph.connectome.FLYNN, sparx.dynamics.Sparse, spectral_radius) | Wang and Chen, arXiv 2607.00025, eq. 1; their LeakyConnectomeRNNCell (models/connectome_rnn_model.py) and spectral rescaling (core/utils.py), ben-gitdev/fly-gym 8d964599 | Against their cell run in PyTorch in float64 (tools/make_flynn_fixtures.py) on a random connectome of 40 neurons, 240 signed synapse counts, five cell classes and unknown, 6 input and 5 output neurons, 15 steps: the output neurons’ activity every step to 1e-12 (observed 7.8e-16) and the gradients of the weights, biases and leak logits to 1e-12 of their largest entries (observed 4.3e-16); reversing the edges, dropping their 0.99 sigmoid + 0.01 or reading their update fraction as the decay fails it. A stream fed in chunks equals one call, and the model trains through dew’s Supervised | Their code’s form of eq. 1, h <- (1 - a) h + a tanh(...) with a per class, is the paper’s with its alpha = 1 - a. The weights start at the synapse counts scaled to the exact spectral radius (observed 0.9 to 1.9e-8 relative; ARPACK agrees with the dense eigenvalues to 2.3e-15 at 2,100 neurons), where theirs scales by a power iteration’s Rayleigh quotient from a random start: on the fixture’s connectome it estimates 20.1 against an exact 48.1, and other starts give 1.8 to 46, so their starting radius depends on the draw. Their readout network and input scales stay outside, as the caller’s layers |
Rate unit (sparx.dynamics.RateCell, nn.Rate) | FLYNN (Wang and Chen, arXiv 2607.00025), equation 1, h_{t+1} = alpha h_t + (1 - alpha) tanh(W h_t + x_t + b); their code was not run | The equation as the paper writes it, in float64 (reference.flynn): through RecurrentCell, within 1.2e-7 (tanh), 6.0e-8 (sigmoid) and 5e-8 relative (ReLU), per-unit leaks and biases (test_a_recurrent_rate_cell_is_flynns_recurrence); as a Network on a sparse 40-neuron graph with sensory input on 5 neurons, through edge and dense delivery, to 3.3e-16 (test_flynn_on_a_sparse_connectome_is_their_recurrence); a step of dt leaks by exp(-dt / tau) | FLYNN trains alpha directly, one per cell class. decay is that alpha per unit of time, raised to dt, so FLYNN’s unit is dt = 1. The recurrent product reaches the unit as its jump, through RecurrentCell or a projection onto a Delta receptor with a delay of one step, which delivers h_t while h_{t+1} is computed. |
Graded-potential neuron (GradedPotential) | The passive membrane’s analytic solution; the graded transmission of Prinz, Bucher and Marder (Nature Neuroscience 2004), s(V) = 1 / (1 + exp((V_th - V) / delta)) | Under current steps toward -80 to -35 mV, the voltage is E_L + I / g_L (1 - exp(-t / tau_m)) to 6.4e-14 mV and the release its sigmoid to 2.8e-15, in float64 (test_a_graded_potential_neuron_is_the_passive_membrane_and_releases_by_its_sigmoid) | The output is the release at the end of the step, held over the next. Prinz et al.’s synaptic activation relaxes with a time constant that depends on the presynaptic voltage; sparx’s Graded synapse fixes it. |
Graded synapse (Graded) | The analytic response of the two-stage linear filter, release to synapse to membrane | A release pulse of 30 ms through Graded(3.0) into a passive membrane, to 4.3e-14 mV in float64 (test_a_graded_synapse_drives_the_membrane_as_the_analytic_two_stage_filter); a network’s graded projection equals the presynaptic population run alone with its weighted release replayed into the synapse, exactly, through edge and dense delivery | The release is held over each step (zero-order hold), so the synapse is exact for the held input and first order in how fast the release changes within a step. A spiking population onto a Graded synapse, or a graded one onto a spiking synapse, is refused. |
Stochastic release (StochasticRelease) | The binomial model of quantal release (del Castillo and Katz 1954) | 200 synapses per neuron releasing with p = 0.3 and a quantum of 2.5 times the weight: the mean and variance of the input are those of 2.5 Binomial(200, 0.3) within 4 standard errors (observed 0.02 and 0.74 for the mean, 2.4 and 0.8 for the variance), through edge and event delivery, and neurons that share every presynaptic spike are uncorrelated (mean pairwise correlation 0.002); with Tsodyks-Markram depression, the input is Binomial(200, p u x) step by step (z-scores with mean 0 and standard deviation 1 within 0.05) (tests/test_signalling.py) | The resources of a depressing synapse deplete by their mean, as in the deterministic model; depletion by the releases that happened (Fuhrmann et al. 2002) is not modeled. Edge and event delivery draw from different streams, so they agree in distribution and differ draw by draw. A dense projection cannot draw per edge and is refused. |
Gap junctions (GapJunction) | The analytic solution of two coupled passive cells; NEST 3.10 gap_junction between hh_psc_alpha_gap neurons with their sodium and potassium conductances at zero (tools/make_nest_gap_fixtures.py) | At 5 and 50 nS, 0.1 ms steps: within 1.9e-4 and 7.1e-3 mV of the analytic solution, falling fourfold each time dt halves (test_gap_junctions_converge_at_second_order_to_two_coupled_cells); NEST with waveform relaxation is within 1.5e-5 mV of it, and sparx is within the same 1.9e-4 and 7.1e-3 mV of NEST (test_gap_junctions_agree_with_nests_waveform_relaxation); a junction between two populations is the same junction within one | NEST iterates waveform relaxation to a tolerance with cubic interpolation of the partners’ voltages; sparx takes one iteration, predicting the partners’ end voltages and interpolating linearly, which is second order and costs a second advance of each coupled population. NEST without waveform relaxation holds the partner’s voltage of the step before and is 0.24 and 2.1 mV off. At 4,000 nS (ten times C / dt) sparx stays bounded, but the fast mode, shorter than a step, is several mV wrong in its transient. A partner that fires is interpolated toward its reset voltage, as NEST interpolates it. |
Neuromodulator (Modulator) | Its equation, dc/dt = -c / tau + release sum_k delta(t - t_k) | The concentration equals the sum of exponentials over the spikes of the named source neurons to 2.3e-7 in float32 at 0.5 ms steps, and a Plasticity rule reads that value after each step (test_a_modulator_follows_its_equation_and_plasticity_reads_it) | One concentration per modulator for the whole network, with no spatial extent or receptor kinetics, released by spikes only. DopamineSTDP reads it. |
Learning rules
Section titled “Learning rules”| Rule | Reference | Checked | Departures and choices |
|---|---|---|---|
e-prop (sparx.learn.eprop) | Bellec et al. 2020; their numerical verifications (numerical_verification_eprop_*.py in IGITUGraz/eligibility_propagation) | Their two identities to 1e-9 relative (tests/test_learn.py): the online gradients equal BPTT’s with the gradient stopped at the recurrent spikes (their stop_z_gradients=True, RecurrentCell(cut_gradient=True) here), and the eligibility traces weighted by the true learning signal equal BPTT’s gradient (their equation 1). Checked for LIF and ALIF with detached resets (ALIF with and without their refractory period of 2 steps, as their verifications run it), LIF with the reset in the gradient, LIF with a decay per neuron, and the current-based Serial(LICell, LIFCell); and at dt = 2 against the same network at dt = 1 | Their code is TensorFlow 1 and does not run here, so the identities, which they use to verify their own code, are the reference. The eligibility traces are derived for any elementwise model by forward-mode derivatives of its step, not hand-coded per model; a leaky readout’s learning signal is handled by filtering the traces with its leak, which keeps the rule online and exact. A state variable whose input enters with constant coefficients shared by all neurons (the membrane under a detached reset) keeps their filtered presynaptic trace z_bar in place of a vector per synapse, and one no gradient reaches (the refractory count) keeps nothing; which is which is read from the jaxpr of the step’s derivative, and an unclear case keeps the full vector. On SHD’s shapes this is 14 times faster than the vector per state variable for ALIF (benchmarks/bench_eprop.py, performance). |
OTTT (sparx.learn.ottt) | Xiao et al. 2022; their OnlineLIFNode and WrapedSNNOp (pkuxmq/OTTT-SNN) | Every weight and bias gradient of a two-hidden-layer network over 12 steps, against their modules run in PyTorch, to 1e-10 relative (tools/make_ottt_fixtures.py) | None found. Their trace starts at the first spike undecayed, which the recursion from zero reproduces. ottt takes the trace’s time constant tau and decays it by exp(-dt / tau); their trace decays by 1 - 1 / tau, which the test passes as the time constant -1 / log(1 - 1 / tau). |
REINFORCE (sparx.learn.reinforce, policy_gradient) | Williams 1992: grad E[R] = E[R grad log P(s)], for a recurrent layer of BernoulliCells | On a layer small enough to enumerate (two neurons with recurrence, three steps, 64 trajectories, both reset rules, float64): every eligibility equals the gradient of its trajectory’s log-probability written independently (observed 8.9e-16), the probabilities sum to 1, the expected eligibility is 0, and the probability-weighted sum of reward times eligibility is the exact gradient of the expected reward (observed 1.7e-16). Sampled: policy_gradient over 20,000 trajectories lies within 5 standard errors of that gradient, with and without a baseline at the expected reward, which cuts its spread 3.9-fold (the test asks for twofold). A hundred steps of ascent teach a layer to fire one neuron and silence another. Dropping the zero reset’s derivative fails the enumeration (tests/test_learn.py) | The derivative of log P(s_t) in the membrane is beta (s_t - p_t), written from the cell’s sigmoid, not differentiated through it: the ratio (s - p) / (p (1 - p)) is 0 / 0 in float32 once p rounds to 1. The weights’ eligibilities only; the decay, threshold and beta get none. |
Reward diffusion and its network (sparx.learn.RNeuralNet, reward_diffusion) | RNeuralNet-Research d4b7803: Global_Refresher, Global_ForwardProcessor, Global_Adjuster, Global_RewardSpreader, Global_Teacher and Global_Renew (Processor.cpp), Soma_t, Neuron_t and Neurite_t (neuron.h, neuron.cpp), NeuralNet_init (Main.cpp); the rule as the owner’s research notes reconstruct it | The original’s classes and functions compiled unchanged (g++ 13.3, -O2) with a driver that runs them single-threaded in the order RNeuralNet describes (tools/make_rneuralnet_fixtures.py, tools/rneuralnet_driver.cpp), on 40 neurons of 4 connections each with delays of 1 to 21 ticks, 3 input and 2 output neurons, for 80 ticks with three rewards: every neuron’s output each tick to 1e-6 (observed 7.2e-7 on outputs up to 9.4, on both sides of the thresholds), the local rewards to 1e-7 (observed 6.0e-8) and the weights after each reward to 1e-7 (observed 0). With each delay at Myelin + 1, without the tick a connection loses when its source sends after its target in the pass, outputs differ by 10. The iterative first-visit spread equals a recursive transcription of Global_RewardSpreader on random cyclic graphs (observed 1.8e-7), and depends on the order of the connections as the original does; every path’s spread equals (I - discount P)^-1 r (observed 4.8e-7) | The original’s threads share the neurons without synchronization, so RNeuralNet fixes one order of its functions, and calls Global_Adjuster, which Main.cpp leaves commented out. NeuralNet_init makes the output neurons’ connections to the feeder last, which in that order splits their sending in two: wire refuses it and random makes the feeder connection with the neuron’s others. random wires every neuron, where an off-by-one leaves the original’s last one unconnected, and draws from numpy, not rand(). The shares are taken less each target’s largest activity, finite where the original’s float exponentials overflow. The input neurons’ local rewards, which Global_Renew never clears and nothing reads, are not compared. |
Fast weights (sparx.dynamics.FastWeights with DecayingHebb, OjaHebb, ModulatedHebb or RetroactiveHebb, on a RecurrentCell; sparx.nn.Recurrent(rule=...)) | Miconi et al. 2018 (differentiable plasticity, eqs. 1 to 3) and 2019 (Backpropamine, eqs. 3 to 5); their networks: NETWORK of simple/simple.py and Network of maze/maze.py with rule="oja" (uber-research/differentiable-plasticity 5bd29a18), Network of simplemaze/maze.py and of maze/batch.py with type="modul", addpw=3 (uber-research/backpropamine 180c9101) | Each of the four against their class run in PyTorch in float64 over 20 to 23 steps (tools/make_miconi_fixtures.py): every step’s activity and trace to 1e-12 (observed at most 1.3e-14 and 4.5e-14, both in the modulated network), the loss to 1e-12 relative, and every parameter’s gradient to 1e-12 of its largest entry (observed at most 1.0e-14). The clip holds part of each bounded trace, and a mutation of each rule fails its case. On their pattern completion, shrunk to two patterns of 8 bits, every rule trained 300 steps gets 4 to 5 % of the zeroed bits wrong, the same network without a trace 22 %; on their full task (1000 bits, five patterns, 2000 episodes; examples/pattern_completion.py) the decaying trace leaves 0.3 % of the zeroed bits wrong and the network without one 50.1 %. A Sparse wiring that lists every pair computes the Dense layer’s activity and traces under every rule (observed 6.3e-7 in float32), and reading a connection’s presynaptic unit for its postsynaptic one fails it | A step of dt keeps (1 - eta) ** dt of a decaying trace or eligibility and adds dt times the rate to the other traces; at dt = 1 both are theirs. maze/batch.py keeps its matrices [post, pre], transposed here. Their tanh unit is RateCell(0.0); the Rate layer also learns a bias per unit, which their networks take from a bias neuron or an input layer. The fan-out of ModulatedHebb is in their code and their appendix, and RetroactiveHebb has none, as maze/batch.py has none. sparx.nn’s neuromodulator starts as their torch.nn.Linear layers do |
Predictive coding and PC-ALM (sparx.learn.PredictiveCoding, residual_mlp, PredictiveCodingObjective) | Seely and Gould 2026 (arXiv 2605.31022), eqs. 8, 11 and 12 and Algorithm 1; their JAX reference, pcalm.inference and pcalm.model (SakanaAI/pc-alm 660747f6) | Against their code in float64 (tools/make_pcalm_fixtures.py) on two of their residual MLPs (tanh, depth 4; ReLU, depth 8 at T = 2L), for PC and for PC-ALM with either weight-credit timing and with two activity steps per multiplier step: every settled activity to 1e-12 (observed 8.9e-16), every multiplier to 1e-12 (observed 1.6e-15), and every weight update to 1e-12 of its layer’s largest entry (observed 4.8e-14). A mutation of the multiplier’s sign, of the credit timing, of the number of cycles, of the state the update reads, or of the per-example step fails its cases. Backpropagation through residual_mlp gives their BP gradient, and the objective hands the rule’s update to dew’s trainer unchanged (observed 0). Their headline cell (Fashion-MNIST, width and depth 32, one epoch) trained through dew at seeds 0, 1 and 2 scores 75.1, 76.1 and 76.5% by PC-ALM, 62.2, 65.5 and 65.6% by PC and 77.8, 76.9 and 76.7% by BP, against their code on the same machine at 77.73, 76.49 and 76.34%, 68.17, 64.49 and 66.54%, and 78.65, 76.85 and 77.31% | Each example descends its own energy at state_lr, the step their code reaches by scaling the batch mean’s gradient by the batch size; the weight update is the energy summed over the batch’s real rows, which dew’s trainer divides by their count, their mean. The weights start from dew’s draw of their distribution, so a run matches theirs in distribution, not draw for draw |
Event-based exact gradients (sparx.learn.events) | Wunderlich and Pehle 2021 (EventProp): the exact gradient of LIF networks with exponential current synapses in continuous time | Gradients by weights and by input spike times against central differences of the exact simulation, to 1e-6 relative in float64 (test_event_gradients_are_the_exact_derivatives_of_spike_times); the event simulation fires as sparx.dynamics’ LIF integrated at 1 us does, spike for spike within 3 us, and a burst of 19 spikes between two inputs fires every spike the grid does, each a threshold crossing of the closed-form membrane to 1e-12 (test_a_strong_pulse_fires_every_spike_until_capacity_runs_out); a two-layer network learns a latency task | The gradient is EventProp’s, obtained differently: each threshold crossing is found by bisection on the membrane’s monotone rising branch and then differentiated by the implicit function theorem (one Newton step with the root held fixed), and autodiff of the event loop does the rest, where EventProp integrates adjoint dynamics backward between the spikes. Both are the exact derivative with spike counts fixed; this one stores each event of the forward pass, EventProp’s adjoint stores only the spike times. Their reference code was not run. |
ANN-to-SNN conversion (sparx.learn.convert) | Rueckauer et al. 2017 (robust normalization, IF with reset by subtraction, analog input, batch norm folding, average pooling, max pooling gated by the most active input); their toolbox, snntoolbox 0.6 | Their rate bound: a neuron under a constant drive a in threshold units fires at clip(a, 0, 1) within 1 / T (T = 10, 100, 1000). A ReLU MLP converted with the 99.9th percentile reaches its accuracy within 2% at 300 steps. A CNN with batch norm, max and average pooling keeps the ReLU network’s predictions on all 500 test images at 300 steps, and hidden rates correlate with normalized activations above 0.99. Folding batch norm reproduces flax’s outputs to float precision. Against snntoolbox run on a Keras CNN (tools/make_snntoolbox_fixtures.py): the folded and normalized weights match to float32 precision, the first layer fires the same spikes, and predictions agree on all 50 test images (output rates correlated at 0.99991). A converted MLP with reset to zero exports to NIR and imports back with the same weights and spikes | Average pooling passes averaged spikes to the next layer; snntoolbox puts IF neurons on it, which differs by up to 1 / T. The max-pool gate counts spikes before the current step; snntoolbox counts the current step too, which lets a pool outfire its most active input (by up to 0.21 on the fixture), so spikes after the first max pool differ (with their rule, the pool and the next layer fire the same spike counts as theirs). Unpadded pooling and 2D NHWC convolutions only; the spiking softmax is not converted. The converted network flattens with sparx.nn.Flatten, channels first; the toolbox’s dense layer reads Keras’s channels-last flatten of a 1 x 1 map, where the two orders agree, and the test reorders its rows explicitly all the same. |
Exchange
Section titled “Exchange”| Format | Reference | Checked | Departures and choices |
|---|---|---|---|
NIR (sparx.nir) | NIR 1.0.8; snnTorch 1.0.0’s export_to_nir and import_nir | Three networks exported by snnTorch (two dense layers; two Conv2d layers with Flatten; RLeaky) import into sparx and fire spike for spike as snnTorch ran them, hidden layers included; exported back, every weight, bias, time constant, threshold, stride, padding, dilation and edge equals snnTorch’s graph; a sparx stack of convolutions, Flatten, Recurrent(LIF) and dense layers round-trips through NIR to identical spikes, and a classifier with an LI readout to identical membranes under both discretizations (tests/test_nir.py) | NIR’s neurons are continuous, and each library picks its step: snnTorch reads them by forward Euler at a fixed 1e-4 s. Sparx states the discretization on both import and export ("exact", its own, or "euler"). NIR is channels first and sparx channels last; weights are transposed and sparx.nn.Flatten orders features C, H, W as PyTorch does. snnTorch’s RLeaky feedback has a bias that sparx’s lacks; it moves into the previous layer’s bias and exports back as a Linear node. snnTorch’s import_nir fails on recurrent graphs, so sparx’s recurrent export is checked against its own import. NIR 1.0.8’s type check rejects non-square and grouped kernels; sparx imports them from an unchecked graph only. snnTorch’s exporter writes scalar parameters with a shape NIR’s type check rejects; the fixtures give per-neuron parameters. Only hard-reset LIF and IF neurons, and LI readouts, so far; an IF node’s r is 1 / dt under either discretization. NIR’s LI is its LIF without a threshold, so it shares the LIF’s discretization, which the snnTorch fixtures check; no library’s own LI export is compared. |
Encoders and losses
Section titled “Encoders and losses”| Function | Reference | Checked | Departures and choices |
|---|---|---|---|
encode.LatencyEncoder | snnTorch spikegen.latency(linear=True, normalize=True, clip=True) | Exact | snnTorch fires values below the threshold at the last step unless clip=True; sparx never fires them. |
encode.DeltaEncoder | snnTorch spikegen.delta(padding=False) | Exact, with and without off-spikes | None. |
losses.per_step_cross_entropy | snnTorch ce_rate_loss | Within 1e-6 relative | None. |
losses.rate_mse | snnTorch mse_count_loss | Within 1e-6 relative, as snnTorch’s value divided by T | snnTorch’s squared error of counts over T grows with T; sparx’s rates do not, and snnTorch rounds its target counts down. |
losses.van_rossum | van Rossum 2001; Elephant 1.2.1 van_rossum_distance | All pairs of six random trains at three time constants, to 1e-10 (tools/make_elephant_fixtures.py) | Elephant’s distance is sqrt(2) times van Rossum’s: it drops the 1/2 of the closed form. Sparx keeps the paper’s definition. The integral is exact for spikes on the grid, the tail after the last step included, not a Riemann sum. |
spiketrains.victor_purpura | Victor and Purpura 1996; Elephant victor_purpura_distance | All pairs, three costs, exactly | None. |
spiketrains.coincidence_factor | Kistler et al. 1997, in Jolivet and Gerstner’s (2004) form; brian2modelfitting 0.4 get_gamma_factor(rate_correction=False) | 21 pairs of trains over 2 s at three precisions: the same coincidences, within 4e-15, and the same Γ where the spike counts are equal, within 6e-17 (tools/make_gamma_fixtures.py) | brian2modelfitting takes the chance rate ν from the reference train. Jolivet and Gerstner take the predicting train’s, under which a Poisson train of any rate scores 0 on average, and sparx follows them. A neuron silent in both trains has no Γ (NaN). |
| Surrogates | The papers’ formulas | Gradients and forward-mode tangents, and each surrogate’s area | None. |
losses.softmax_sum_cross_entropy, readout="softmax_sum" | SNN-delays’ calc_loss with loss='sum' and its prediction | Their formula in float64, within 1e-6 relative, and through their whole network (above) | None: their summed probabilities go through torch’s CrossEntropyLoss, which takes a log-softmax of them again, and sparx does the same. |
datasets.bin_events(binning="events"), shd(binning="events") | SpikingJelly’s integrate_events_by_fixed_duration_shd in the 2023 tree SNN-delays runs on | Exact on a recording with silences (test_event_binning_reproduces_snn_delays_frames), and on all 10,420 SHD recordings against the frames their running code cached, at 10 ms with channels summed in fives | Their frames open at events: each takes the events within 10 ms of the first one not yet counted, so silences longer than a step vanish and recordings shorten (at most 124 steps where the grid needs 137). Later SpikingJelly releases bin on a grid from the first spike. Times are scaled to ms in the file’s float16, as theirs are. |
Training
Section titled “Training”| Piece | Reference | Checked | Departures and choices |
|---|---|---|---|
dew’s OneCycle, Cosine with no warmup and Exponential, stepped once an epoch (every) | SNN-delays’ schedulers: torch’s OneCycleLR(max_lr=5e-3) for the weights and its Adam momentum cycle, CosineAnnealingLR for the positions, decrease_sig for the width, each stepped at the end of an epoch | All 150 epochs of their SHD configuration, recorded from their code: learning rates within 1e-5 relative (the cosine within 1e-4), momentum within 1e-6, width within 1e-5 (test_dews_schedules_stepped_once_an_epoch_are_snn_delays_torch_schedulers) | None. |
dew’s OptimConfig(optimizer="adam") with a ParamGroup each for SNN-delays’ delays, weights and norms | torch’s Adam(weight_decay=...), whose decay is an L2 term added to the gradient, with OneCycleLR setting its lr and betas[0] before each step, and DCLS’s clamp of the positions | Dew’s own tests, against torch’s update (tests/test_optim_schedules.py in dew); sparx checks that the groups find a DelayedDense’s delays and keep them within their bounds (test_dews_parameter_groups_find_the_delays_and_keep_them_within_bounds) | Groups are matched by path patterns; torch takes lists of tensors. |
examples/train_shd.py --recipe snn-delays | Their best_config_SHD.py and training loop | The network (214,016 parameters, their count), loss, schedules and binning, as above. The first 2 epochs of the 20-epoch schedule on SHD with a 10% holdout: test accuracy 8.7% then 19.8% (validation 9.4%, 17.9%), where their code on the same schedule and all of the training set scores 8.1% then 23.5%; 7.5 to 14 s a training step of 256 on a shared 4-core CPU | Their batches are padded to the batch’s longest recording (about 105 steps on average); sparx pads every recording to 124, so the summed softmax and the batch statistics also see the extra silent steps. dew drops an epoch’s last partial batch (31 steps of 256, where they take 32 with a last batch of 220). Their test accuracy is the mean of per-batch accuracies; sparx’s is over all 2264 recordings. They select epochs on the test set and report the best test accuracy; the example holds out 10% of the training set by default and reports the test accuracy at the best validation epoch beside theirs. Random streams (initialization, shuffling, dropout) differ. |
- Shiu et al.’s model on the male CNS is calibrated, not validated: a published simulation on this connectome, or more behaviours than sugar to MN9, would test it. The left and right sugar neurons drive MN9 unequally.