Skip to content
GitHub

Fidelity ledger

For each model sparx implements: the references it is checked against, every difference found between a reference and the science or another reference, what sparx does, and the test that pins it. design.md describes the tiers. A row is added before a model ships, not after.

Reference versions: SpikingJelly at commit c6cb8e46 (2026-10-03), snnTorch 1.0.0, DCLS 0.1.1 (Hammouamri et al.’s layer), IGITUGraz/eligibility_propagation at its default branch (Bellec et al.), Thvnvtos/SNN-delays at commit d169b4e3 on the 2023 SpikingJelly tree it was written against (Hammouamri et al.), all with torch 2.14.1+cpu; NEST 3.10.0 and Brian2 2.10.1 for the biophysical models. The fixture tools under tools/ record the versions in each fixture.

ModelReferenceCheckedDepartures and choices
LIF (sparx.dynamics.LIFCell, nn.LIF)SpikingJelly LIFNode(decay_input=False); snnTorch Leaky(reset_delay=False)Spikes exact and gradients within 3.6e-7 of both, soft, hard and detached resets (tests/test_reference.py)Sparx’s decay is exp(-1/tau), the exact decay of the leak over a step; SpikingJelly’s is 1 - 1/tau, the Euler step. The model takes the decay itself, so either convention is a value, and the parity tests pass SpikingJelly’s.
LIF reset timingsnnTorch’s default reset_delay=TruePinned as different (test_snntorchs_default_delayed_reset_is_a_different_model)snnTorch’s default subtracts the threshold one step after the spike, without decaying it. That is no discretization of the continuous LIF, whose reset happens at the spike and then decays. Sparx resets on the spiking step.
LIF after an overshootsnnTorch reset_delay=FalsePinned (test_snntorch_loses_spikes_after_an_overshoot_where_sparx_fires)When a soft reset leaves the membrane at or above threshold, snnTorch subtracts the next reset before testing for a spike, so the neuron needs twice the threshold to fire again and loses spikes. Every neuron where the two diverge does so on the step after its own overshoot. Sparx and SpikingJelly fire whenever the membrane is at threshold, removing one threshold per spike.
LIF spike conditionsnnTorch fires on v - threshold > 0; SpikingJelly and sparx on >= 0Fixtures keep every membrane 1e-4 from thresholdDiffers only at exact equality.
Detached resetsnnTorch detaches its delayed reset and keeps the gradient of its immediate oneGradients match with detach_reset=FalseSparx’s detach_reset is SpikingJelly’s option of the same name.
Current-based LIF (Serial(LICell, LIFCell), nn.Synaptic)snnTorch Synaptic(reset_delay=False)Spikes exact and gradients matched, soft and hard resetSame reset timing and overshoot departures as LIF. The synaptic current is an LICell whose membrane feeds the LIFCell each step, the same operations in the same order as snnTorch’s single cell.
Adaptive LIF (sparx.dynamics.ALIFCell)Bellec et al. 2020 equations; their code, Figure_4_and_5_ATARI/alif_eligibility_propagation.pySpike for spike against a transcription of their cell with a reset of decay * threshold (test_alif_is_bellecs_model_with_its_reset_decayed)A spike subtracts the baseline threshold, as theirs does (sparx subtracted the adaptive threshold before this audit). Their reset lands one step later, undecayed, so their model with reset decay * threshold is sparx’s. Their repository has a second variant, LightALIF, that scales input by 1 - decay and adaptation by 1 - adapt_decay; sparx follows the paper’s equations, which the ATARI cell implements. Refractoriness is their counter (refractory, their n_refractory, a duration of round(refractory / dt) steps): spike for spike against a transcription of it (test_alif_refractoriness_is_bellecs_counter), and off by default. Their pseudo-derivative is Triangle(width=threshold, scale=0.3 / threshold) on v - A; sparx’s default surrogate is ATan for every neuron.
Escape-noise LIF (sparx.dynamics.BernoulliCell)Its equations, the discrete-time escape noise of Pfister et al. 2006: the LIF membrane, firing with probability sigmoid(beta (v - threshold)) per stepSpike for spike against a float64 loop given the same uniform draws, each reset rule, probabilities within 1.9e-7 (test_bernoulli_matches_the_reference_loop); noise 1 - s replays the spikes s; a run in two chunks equals oneThe noise is an input (SynapticInput.noise), drawn by the caller, so the cell is a pure function and a trajectory can be forced. p is per step whatever dt is. The spike is a sample and passes no gradient.
RNeuralNet’s neuron (sparx.dynamics.PulseCell)Soma_t::ActivationFunction of RNeuralNet-Research (AshishKumar4/RNeuralNet-Research d4b7803), A_CONST 1Against the original program (RNeuralNet under learning rules) and its formula in float64 (test_a_pulse_cell_is_rneuralnets_activation, observed 7.0e-8)The original’s sum empties once the neuron has sent on all its connections; in the order RNeuralNet runs it, that is every tick, so the cell keeps only a jump that lands after its step. exp(min(s - threshold, 0)) - 1 keeps the unused branch finite and equals the original’s below the threshold.
Izhikevich (sparx.dynamics.Izhikevich, scheme="published", order="izhikevich", the defaults; nn.Dynamics(Izhikevich()))Izhikevich 2003, the paper’s MATLAB codeSpike for spike in float64 against a transcription of the code (test_izhikevich_matches_his_published_loop_spike_for_spike_in_float64); regular spiking under input 10 adapts, then fires every 47 to 62 msEach 1 ms step takes two half-steps of v and one step of u, as the code does (sparx took 0.5 ms Euler steps of both before this audit). The firing patterns the model is known for are those of this scheme; an integration closer to the ODE is a separate, named model (design.md section 4.2). The peak is not clipped at 30 mV, as in the 2003 code.
Izhikevich numericsXLAMeasuredCompiled XLA rounds fused arithmetic differently from NumPy in the last bit (22% of elements of one step). The quadratic membrane amplifies that chaotically, so exact comparisons run op by op (jax.disable_jit).
PSN, MaskedPSN, SlidingPSNSpikingJelly neuron/psn.pySpikes exact, gradients within 1.2e-6None found. MaskedPSN’s lambda_ is the call argument masking.
Biophysical LIF (sparx.dynamics.LeakyIntegrateAndFire)Its equation’s exact solution over a step; NEST iaf_psc_exp, iaf_psc_alpha, iaf_psc_delta; Brian2 exactGround truth in float64 (tests/test_dynamics.py): one 5 ms step equals 500 steps of 0.01 ms to 1e-12, conductance relaxation and f-I periods are the analytic ones. Against NEST (tests/test_simulators.py): spike for spike and voltages within 1e-11 mV over 300 ms of random weighted input, for exponential, alpha and delta synapsesNone for current input. Refractoriness holds round(t_ref / dt) steps after the spiking step, as NEST does.
RefractorinessBrian2Spike for spike and within 1e-9 mV once sparx’s t_ref is one step shorter (test_brian2_*)Brian2 counts t_ref from the start of the step in which the neuron crossed, the time it stamps on the spike, so it holds one step less than NEST. The crossing happened within that step, so Brian2’s period is up to a step shorter than t_ref and NEST’s up to a step longer; sparx follows NEST, which never undercuts the stated period.
Conductance-based LIFNEST iaf_cond_exp, iaf_cond_alpha, iaf_cond_beta (adaptive RK45); Brian2 exponential_euler; a float64 RK4 at dt / 10NEST is within 1.5e-9 mV of the RK4 truth. Sparx’s default (hold="mean") is within 2e-3 mV of NEST and fires with it spike for spike until a crossing within that margin; its error falls fourfold when dt halves (test_conductance_hold_converges_at_second_order_to_the_rk4_truth). hold="start" is Brian2’s scheme: spike for spike and within 1e-9 mVWithin a step the membrane sees a held conductance. Brian2 holds its value at the start of the step, first order, and fires several steps late against NEST over 300 ms. Sparx holds its exact average over the step, which integrates the leak’s decay exactly and is second order, at the same cost. NEST’s adaptive integration is more accurate still and does not vectorize over a population.
AdEx (sparx.dynamics.AdEx)Brette and Gerstner 2005; NEST aeif_psc_exp, aeif_cond_exp (adaptive RK45); Naud et al. 2008’s firing patternsEach of Naud’s eight parameter sets for 500 ms: the same spike count as NEST and every spike within 0.8 ms (ten substeps), within 0.1 ms (a hundred); the irregular set is chaotic, and its spike count and irregularity match. The published patterns hold (tonic, adapting, initial burst, regular bursting, delayed accelerating, delayed regular bursting, transient). With synaptic input, spike for spike within 0.2 ms (current) and 0.5 ms (conductance)NEST’s right-hand side reads the voltage capped at v_peak, and v_reset while refractory; sparx reads it the same way. NEST resets inside its adaptive steps and integrates the rest of the step from the reset, and sparx does so at the substep that crosses: without that, spikes drift 5 ms late over 500 ms whatever the substep count. The upswing is stiff, so the default is ten RK4 substeps per step; an exponential Rosenbrock step was less accurate than RK4 at every count tried.
Izhikevich in physical time (sparx.dynamics.Izhikevich)Izhikevich 2003; NEST izhikevich, both schemesWith order="nest", op by op in float64, NEST’s run to the last bit for 300 ms at 1 ms (test_izhikevich_published_scheme_is_nests_to_the_last_bit); compiled at 0.1 ms, each of the paper’s seven classes fires as many spikes as NEST, every one within 1 ms, under both schemes; delta input spike for spike within 1e-9 mVNEST starts u at -13 whatever b is; sparx starts it at b * v, as Izhikevich’s code does, and the fixtures set NEST’s U_m so. NEST computes (0.04 v) v, Izhikevich’s code 0.04 (v^2); the two differ in the last bit, and the field order names which one a model computes: "izhikevich" by default, "nest" in the NEST tests. NEST adds a delta input to the voltage under Euler and as a one-step current of the weight under the published scheme; sparx’s jump is always a voltage jump, and the published convention is the weight passed as current.
Izhikevich (2004)‘s twenty firing patterns (izhikevich_2004, scheme="semi_implicit")His figure1.m (archived copy), run unchanged but for its plotting in GNU Octave 8.4 (tools/make_izhikevich_2004_fixtures.py)Each of the twenty panels with his current, step and start: every spike on his step, and the voltage between spikes his to a relative 1e-9 in float64, compiled (test_izhikevich_2004_patterns_are_his_codes_runs)His code takes one Euler step of v and then one of u from the new v, which is neither the 2003 code’s scheme nor NEST’s, so it is a third, named scheme. Two panels change the equations: class 1 excitability and the integrator use 4.1 v + 108, accommodation du/dt = a b (v + 65); these are fields (quadratic, v_u, u_decay), not hidden cases. In the class 2 panel the voltage differs by up to 0.012 mV and no spike moves: Octave’s V^2 is libm’s pow, which rounds differently from v * v in about one value in a thousand, and that panel’s slow ramp through the bifurcation amplifies it.
Hodgkin-Huxley (sparx.dynamics.HodgkinHuxley)Hodgkin and Huxley 1952 in NEST hh_psc_alpha’s statement (rates, 100 pF, peak detection, 2 ms refractoriness); Brian2 exponential_eulerFour currents from rheobase to 4 nA for 300 ms: every spike within one step of NEST’s (default), spike for spike and within 0.02 mV with rk4 at 0.025 ms; with alpha synapses under inhibition down to -128 mV, spike for spike. exponential_euler is Brian2’s within 1e-10 mV for 10 ms. The default scheme converges at second order (test_hodgkin_huxley_strang_is_second_order)NEST integrates by adaptive RK45. Fixed-step RK4 is more accurate per cost but diverges once inhibition makes the gates stiff (beta_m reaches 130/ms at -128 mV), and in a network that is not rare; sparx’s default is a Strang splitting of exact gate and voltage solves, stable at any step. Brian2’s exponential_euler (Rush-Larsen gates, exact voltage, first order) fires up to 1.6 ms off NEST at 0.1 ms. A spike is the voltage’s peak, reported at the step where it has begun to fall, as NEST does.
NMDA magnesium block (MgBlock)Jahr and Stevens 1990The formula (test_mg_block_is_jahr_and_stevens); applied at the voltage at the start of the stepHeld over the step like the conductance it scales.
LayerReferenceCheckedDepartures and choices
Exponential, Alpha, BiExponential, Delta synapsesNEST’s iaf_psc_* and iaf_cond_* kinetics; peaks normalized to the weight, as NEST doesThrough the neuron tests above, and analytically: decay, peak value and time (test_dynamics.py). The response of the membrane to each waveform is the exact integral, checked against adaptive quadrature to 1e-12 at equal, close and distant time constantsSpikes land at the end of their arrival step and shape the membrane from the next, NEST’s order (an event stamped T lands at the end of the step ending at T). Brian2’s on_pre with no delay lands at the same point. A synapse declares where its arrivals land (lands): Delta before the threshold test, as NEST’s iaf_psc_delta input, or after it with after_threshold=True, as Brian2’s on_pre="v += w", through the neuron model’s after_threshold, which drops the jump where a reset sets the membrane and keeps it under a subtracting reset or none.
DelayedDenseDCLS Dcls1d (version gauss in training, max with rounded positions at evaluation), as SNN-delays uses itOutputs within 2.4e-7, gradients (input, weight, delay) within 1.4e-6, both modes (test_delayed_dense_matches_dcls_delays)Sparx’s sigma is DCLS’s effective width, abs(SIG) + 0.27, and its delay is K - 1 - (P + K // 2). SNN-delays also pads the input on the right by (K - 1) // 2, which lengthens the output so delayed spikes after the input’s end reach the readout; the layer keeps the length (it is causal and streams exactly), and SpikingMLP(extend=True) appends those zeros to each delayed synapse’s input, their setting. DCLS clamps its positions after every update and has no clip in its kernel, so a delay at an end of the range gets the kernel’s full gradient; sparx’s kernel clips the center with the gradient passed straight through (a plain jnp.clip halves it at the end, which the network test below caught), and dew’s ParamGroup(bounds=...) clamps after each update. Their width decays exponentially to SIG = 0.23 (effective 0.5) over the first quarter of training, dew’s Exponential(init=12, end=0.23, decay_steps=epochs // 4, offset=0.27) for their 25-step kernel; the alif example uses a linear schedule to 0.5.
Delays per connection (sparx.dynamics.Sparse with delay and longest_delay, in a RecurrentCell)Messages queued on each connection and weighted when sent, as RNeuralNet-Research’s Neurite_t queues themAgainst a float64 loop that queues each message for its delay (test_a_delayed_sparse_recurrence_is_the_queued_loop, observed 4.8e-7 on outputs up to 7.4; without the delays the outputs differ by more than 0.1), and through RNeuralNet against the original program; a stream fed in chunks equals one run; fast weights on delayed connections against a float64 loop (test_fast_weights_on_delayed_connections_pair_what_each_one_delivers)A message keeps the weight it left with, so a change reaches only later messages. On a delayed wiring, fast weights pair each connection’s new output with what it delivers this step, its source’s output delay steps before (RecurrentState.history), and a message carries the fast weights of the step that sent it.
SEW ResNet blocksSpikingJelly model/sew_resnet.py BasicBlockSpike for spike in float64 for ADD, AND and IAND, with and without the downsampling shortcut (test_sew_block_matches_spikingjelly)None found. The small stem (one 3x3 convolution, no pooling) is sparx’s addition for 32x32 inputs; the imagenet stem is SpikingJelly’s.
TsodyksMarkramTsodyks and Markram 1997; NEST tsodyks2_synapseThe weight transmitted at every presynaptic spike, depressing and facilitating parameter sets, four synapses for 2 s: within 1e-10 relative (tests/test_plasticity.py)NEST transmits its initial w u x at the first spike and its initial u is 0.5 whatever U is; a synapse at rest has u = U and transmits w U, which sparx does, and the fixtures set NEST’s u so.
PairSTDPGuetig et al. 2003 (additive at mu = 0, Song et al. 2000; multiplicative at mu = 1); NEST stdp_synapseEvery transmitted weight within 1e-10 relative, additive and multiplicative, with a 1 ms dendritic delayNone. NEST evaluates the rule at presynaptic spikes from the postsynaptic history; sparx applies the same updates as the spikes happen, in the same order, including NEST’s for coincident spikes (a spike pairs only with strictly earlier spikes of the other side, and potentiation precedes depression).
TripletSTDPPfister and Gerstner 2006, all-to-all; NEST stdp_triplet_synapseEvery transmitted weight within 1e-10 relativeNone. The defaults are the paper’s visual cortex fit; NEST keeps the postsynaptic time constants on the neuron, with defaults of 20 and 110 ms rather than the fit’s 33.7 and 125 ms.
DopamineSTDPIzhikevich 2007; NEST 3.10 stdp_dopamine_synapse (Potjans et al. 2010) with a volume_transmitterEvery transmitted weight within 1e-10 relative (observed 2.1e-12 to 4.7e-12) for four synapses over 2 s with a dopamine neuron arriving 1 ms after it fires: NEST’s defaults, which sink to Wmin; a baseline b = 0.01 under which n - b changes sign between dopamine spikes, unclipped; and bounds reached on both sides (test_dopamine_stdp_is_nests). In a Network, the weights learned equal the rule run alone on the network’s own spikes and dopamine trace (test_dopamine_stdp_in_a_network_is_the_rule_on_the_networks_own_spikes). Replacing the step’s exact integral by c (n - b) dt fails the NEST test.NEST integrates between events and clips at each; sparx integrates the same exact solution step by step and clips after every step. They differ only when n - b changes sign between dopamine releases and the weight crosses a bound in that interval. The rule’s tau_n is a Python number the network checks against its modulator’s tau.
NetworkReferenceCheckedDepartures and choices
Recurrent Network (sparx.graph)NEST, the same edges, weights and per-edge delays (1 to 3 ms)60 neurons, 10% connectivity, mixed-sign weights, per-neuron currents, 500 ms: spike for spike with iaf_psc_exp and iaf_psc_delta networks; with iaf_cond_exp, spike for spike until a held-conductance difference of 1e-3 mV moves one spike by one step (after 50 ms), and total activity within 5% (tests/test_graph.py)A spike sent in step m over D steps lands at the end of step m + D, NEST’s and Brian2’s timing for a delay of D dt. Kinetic synapses accept D = 0 (Brian2’s default); delta synapses need D >= 1. A delay, duration or chunk that is not a whole number of steps raises (sparx.graph.network.whole_steps). Rounding it would move it by up to half a step, in a direction set by floating-point error.
Brunel (2000), sparx.graph.models.brunelNEST’s brunel_delta_nest.py, the same parameters and connection ruleAt 2,500 neurons, in each of Brunel’s four regimes, the excitatory rate, mean interspike CV and population Fano factor of one sparx run fall within NEST’s spread over eight seeds (test_brunel_regimes_match_nests_statistics)At this size (in-degree 200) the regimes are not those of Brunel’s Figure 8, which needs his in-degree of 1,000; the parity holds either way, and NEST at full size gives the same. The slow synchronous regime varies by 25% between seeds in both simulators, so three seeds were not enough to compare it.
Potjans and Diesmann (2014), sparx.graph.models.microcircuitNEST 3.10 running the PyNEST implementation of INM-6/microcircuit-PD14-model (f79f8ac) unchanged, and running the network sparx draws, edge by edge (tools/make_microcircuit_fixtures.py)At a fifth of the neurons and inputs (15,435 neurons, 12M synapses): the neurons, synapse counts and currents the reference derives, exactly, and weights and delays from its distributions; on the network sparx draws, NEST and sparx in float64 spike alike for 300 ms (12,689 spikes); over the reference’s own draws at 15 seeds, each population’s mean rate within 4 standard deviations of NEST’s, and the Kolmogorov-Smirnov distance of its rates, ISI CVs and pairwise correlations from NEST’s within the distances between NEST’s seeds, the reference’s own verification (Dasbach et al. 2021) (tests/test_microcircuit.py)Weights and delays are drawn again where NEST’s redraw draws them again, as the reference’s code does; its model description clips them instead. At a tenth of the neurons and inputs the currents that make up for the lost inputs leave every population but L6I below threshold, and the network falls silent in NEST and sparx alike (L6I at 3.43 Hz in both), so the comparison runs at the reference’s own scale for its verification data. Over one second, layer 2/3’s excitatory rate varies between NEST’s seeds from 0.55 to 0.64 Hz. In float32 the runs part after about 140 ms, as rounding steers a chaotic network.
CUBA and COBA, sparx.graph.models.cuba, cobaBrette et al. 2007’s benchmarks in Brian2: examples/CUBA.py, and its conductance-based counterpart4,000 neurons, 500 ms: both populations’ rates, the excitatory CV and Fano factor of one sparx run fall within Brian2’s spread over four seeds (test_brette_benchmarks_match_brian2s_statistics)Brian2 counts refractoriness one step shorter (see Neurons). Sparx holds conductances at their step average, Brian2’s exponential_euler at the step start. Both act without delay, which sparx’s kinetic synapses take as Brian2 does.
Shiu et al. (2024), sparx.graph.connectome.shiu2024Their Brian2 model (github.com/philshiu/Drosophila_brain_model), and their published runs on FlyWire v630Their neuron model on a small random graph: spike for spike with their Brian2 statement, through edge and event delivery (test_shius_neuron_model_fires_with_brian2_spike_for_spike). The whole brain (127,400 neurons, 14.7M connections), 21 sugar-sensing neurons activated, 10 trials against their 30: at 100 Hz, total spikes 9,675 vs 9,636, rates of the 347 neurons above 1 Hz correlated at 0.9989, MN9 67.1 Hz vs 67.0 +- 6.6; at 200 Hz, 16,909 vs 17,052, 0.9996 over 391 neurons, MN9 92.7 vs 93.3 +- 3.2 (test_shiu2024_reproduces_their_published_sugar_activation, run when their data is present)Their model has quirks sparx reproduces as options, off elsewhere: a spike clears the synaptic variable g, and Brian2’s (unless refractory) on g both freezes it and drops input arriving while refractory; their stimulus lands after the threshold test (Delta(after_threshold=True)). freeze_synapses reads refractoriness from the neuron model (is_refractory), so it applies to any model; for one without refractoriness it drops only the arrivals of the step in which the neuron fired. Their published sugarR run stimulates at 200 Hz, as their notebook says, though their model.py now defaults to 150 Hz.
Male CNS v0.9, Connectome.from_malecnsJanelia’s release tablesThe reader, on synthetic tables (test_malecns_reader_keeps_neurons_and_signs_them_by_transmitter): 165,899 neurons (Traced and Anchor), 25.6M connections, 124M synapses from the v0.9 releaseShiu et al.’s model has no published run on this connectome, so it is checked against their FlyWire results. Their w_syn (0.275 mV) was fit to FlyWire v630’s synapse counts. The male CNS counts more synapses per neuron: a median of 347 input synapses against 206, and a mean of 748 against 414. At their weight the model runs away: the right sugar neurons (types LB3b, LB3c, LB3d and LB4b, the types of their 21) at 100 Hz recruit 18,249 neurons and 451,000 spikes a second. matched_w_syn scales the weight so the median neuron receives FlyWire’s input, which gives 0.163 mV. At that weight the same stimulus recruits 670 neurons and 11,500 spikes a second, against FlyWire’s 400 and 9,600, and drives MN9 (body 10331) at 81 Hz, against FlyWire’s 67 (test_shiu2024_on_the_male_cns_activates_mn9_from_sugar_neurons_at_the_matched_weight, which runs when the release tables are present). This is a calibration, not a reproduction: one scalar matched on one statistic, checked against one behaviour. The left sugar neurons drive MN9 at only 8 Hz, so the two sides differ, and FlyWire’s side labels are not checked against the male CNS’s here.
Simulation over devices (simulate(mesh=MeshSpec(...)))One deviceOn 4 simulated CPU devices of dew’s mesh: NEST’s 60-neuron network with its neurons partitioned over the data axis (15 per device) fires spike for spike as on one device over 300 ms; 8 trials spread over the data axis equal 8 trials on one; with MeshSpec(fsdp=2) the trials split over data and each trial’s neurons over fsdp, with the same spikes (tests/test_graph_distributed.py)The step’s exchange of spikes between devices is left to the compiler’s partitioning of the gathers; a hand-written exchange of spike bitmasks (design.md 6.3) waits for measurements on real multi-device hardware. A population whose size the axis does not divide stays whole on every device, as dew’s logical_spec places any array.
Resumed simulation (simulate(checkpoints=...))The unbroken runA run stopped after 20 ms and resumed from its dew checkpoint to 50 ms gives the unbroken run’s spikes from 20 ms on and its final state, exactly (test_a_run_resumed_from_its_checkpoint_is_the_unbroken_run)The checkpoint holds the state collection and the key, not the connectome or params, which the resumed run takes from the variables it is given. Records start at the checkpoint, and an accumulating monitor sums from there.
Gradients through a NetworkCentral differences in float64; the same synapses delivered as an edge listA network whose driven neurons fire, feeding neurons held below threshold through two trainable projections: the gradient of their membranes’ squared error equals central differences within 2.6e-9 relative (test_gradients_through_a_network_are_exact_where_no_membrane_nears_threshold). A trainable projection feeding a population with a recurrence delivered by events, with one delay or a delay per edge: the gradient equals the one through the same recurrence as an edge list within 8e-15 of the largest (test_gradients_through_event_delivery_are_the_edge_lists)Event delivery is a while_loop, which has no reverse mode; it is differentiated as its edge list, so its backward pass visits every edge at every step. Stochastic release has no gradient. Where neurons fire the gradient is the surrogate’s, and detach_reset=True stops it through the reset.
Plastic and short-term projections——A plastic weight is read when a spike arrives; NEST reads it when the spike is sent, so a weight change while a spike is in flight (within the delay) is applied to it in sparx and not in NEST. Short-term release is computed when the spike is sent and travels with it, as in NEST.
SNN-delays (Hammouamri et al., ICLR 2024), SpikingMLP with every synapse delayedTheir SnnDelays on DCLS 0.1.1 and the 2023 SpikingJelly their code runs on (tools/make_snn_delays_fixtures.py), a small copy of their SHD configuration: two hidden layers, delays of 0 to 4 steps on every synapse, the right padding, batch norm, decay_input=False LIF with hard reset and detach_reset, ATan(5), no bias, the non-spiking LIF readout and loss='sum'A training step (Gaussian kernels, batch statistics): outputs and loss within 1.2e-7, every gradient (input, weights, delays, batch norm) within 5.7e-7, running means within 1e-5. An evaluation as their eval_model runs it (rounded delays, running statistics): outputs within 2.4e-7 (test_a_fully_delayed_network_trains_as_snn_delays, ..._evaluates_as_snn_delays)torch’s running variance is the unbiased one and flax’s the biased one, a factor n / (n - 1) for n = steps x batch, which the test applies. Their LIF keeps 1 - 1 / tau of the membrane a step; sparx’s LIF keeps exp(-1 / tau'), so the recipe passes the tau' that gives their decay. Their dropout holds one mask over a recording (SpikingJelly’s multi-step Dropout), dropout_mask="sequence"; the fixture runs without dropout, since a random mask cannot be matched.
ModelReferenceCheckedDepartures and choices
A model’s output (Output(value, offset), graded)Every reference aboveEvery parity test in this ledger passes unchanged after spikes became Output.value; Shiu et al.’s whole-brain model fires 10,148 spikes in 1 s of sugar stimulation before and after, at the same 1.9 ms per step on CPULICell used to report its membrane as a spike; it is a graded model now, and sparx.nn.LI returns the same values. Event delivery reads spikes only and refuses a graded population.
FLYNN (sparx.graph.connectome.FLYNN, sparx.dynamics.Sparse, spectral_radius)Wang and Chen, arXiv 2607.00025, eq. 1; their LeakyConnectomeRNNCell (models/connectome_rnn_model.py) and spectral rescaling (core/utils.py), ben-gitdev/fly-gym 8d964599Against their cell run in PyTorch in float64 (tools/make_flynn_fixtures.py) on a random connectome of 40 neurons, 240 signed synapse counts, five cell classes and unknown, 6 input and 5 output neurons, 15 steps: the output neurons’ activity every step to 1e-12 (observed 7.8e-16) and the gradients of the weights, biases and leak logits to 1e-12 of their largest entries (observed 4.3e-16); reversing the edges, dropping their 0.99 sigmoid + 0.01 or reading their update fraction as the decay fails it. A stream fed in chunks equals one call, and the model trains through dew’s SupervisedTheir code’s form of eq. 1, h <- (1 - a) h + a tanh(...) with a per class, is the paper’s with its alpha = 1 - a. The weights start at the synapse counts scaled to the exact spectral radius (observed 0.9 to 1.9e-8 relative; ARPACK agrees with the dense eigenvalues to 2.3e-15 at 2,100 neurons), where theirs scales by a power iteration’s Rayleigh quotient from a random start: on the fixture’s connectome it estimates 20.1 against an exact 48.1, and other starts give 1.8 to 46, so their starting radius depends on the draw. Their readout network and input scales stay outside, as the caller’s layers
Rate unit (sparx.dynamics.RateCell, nn.Rate)FLYNN (Wang and Chen, arXiv 2607.00025), equation 1, h_{t+1} = alpha h_t + (1 - alpha) tanh(W h_t + x_t + b); their code was not runThe equation as the paper writes it, in float64 (reference.flynn): through RecurrentCell, within 1.2e-7 (tanh), 6.0e-8 (sigmoid) and 5e-8 relative (ReLU), per-unit leaks and biases (test_a_recurrent_rate_cell_is_flynns_recurrence); as a Network on a sparse 40-neuron graph with sensory input on 5 neurons, through edge and dense delivery, to 3.3e-16 (test_flynn_on_a_sparse_connectome_is_their_recurrence); a step of dt leaks by exp(-dt / tau)FLYNN trains alpha directly, one per cell class. decay is that alpha per unit of time, raised to dt, so FLYNN’s unit is dt = 1. The recurrent product reaches the unit as its jump, through RecurrentCell or a projection onto a Delta receptor with a delay of one step, which delivers h_t while h_{t+1} is computed.
Graded-potential neuron (GradedPotential)The passive membrane’s analytic solution; the graded transmission of Prinz, Bucher and Marder (Nature Neuroscience 2004), s(V) = 1 / (1 + exp((V_th - V) / delta))Under current steps toward -80 to -35 mV, the voltage is E_L + I / g_L (1 - exp(-t / tau_m)) to 6.4e-14 mV and the release its sigmoid to 2.8e-15, in float64 (test_a_graded_potential_neuron_is_the_passive_membrane_and_releases_by_its_sigmoid)The output is the release at the end of the step, held over the next. Prinz et al.’s synaptic activation relaxes with a time constant that depends on the presynaptic voltage; sparx’s Graded synapse fixes it.
Graded synapse (Graded)The analytic response of the two-stage linear filter, release to synapse to membraneA release pulse of 30 ms through Graded(3.0) into a passive membrane, to 4.3e-14 mV in float64 (test_a_graded_synapse_drives_the_membrane_as_the_analytic_two_stage_filter); a network’s graded projection equals the presynaptic population run alone with its weighted release replayed into the synapse, exactly, through edge and dense deliveryThe release is held over each step (zero-order hold), so the synapse is exact for the held input and first order in how fast the release changes within a step. A spiking population onto a Graded synapse, or a graded one onto a spiking synapse, is refused.
Stochastic release (StochasticRelease)The binomial model of quantal release (del Castillo and Katz 1954)200 synapses per neuron releasing with p = 0.3 and a quantum of 2.5 times the weight: the mean and variance of the input are those of 2.5 Binomial(200, 0.3) within 4 standard errors (observed 0.02 and 0.74 for the mean, 2.4 and 0.8 for the variance), through edge and event delivery, and neurons that share every presynaptic spike are uncorrelated (mean pairwise correlation 0.002); with Tsodyks-Markram depression, the input is Binomial(200, p u x) step by step (z-scores with mean 0 and standard deviation 1 within 0.05) (tests/test_signalling.py)The resources of a depressing synapse deplete by their mean, as in the deterministic model; depletion by the releases that happened (Fuhrmann et al. 2002) is not modeled. Edge and event delivery draw from different streams, so they agree in distribution and differ draw by draw. A dense projection cannot draw per edge and is refused.
Gap junctions (GapJunction)The analytic solution of two coupled passive cells; NEST 3.10 gap_junction between hh_psc_alpha_gap neurons with their sodium and potassium conductances at zero (tools/make_nest_gap_fixtures.py)At 5 and 50 nS, 0.1 ms steps: within 1.9e-4 and 7.1e-3 mV of the analytic solution, falling fourfold each time dt halves (test_gap_junctions_converge_at_second_order_to_two_coupled_cells); NEST with waveform relaxation is within 1.5e-5 mV of it, and sparx is within the same 1.9e-4 and 7.1e-3 mV of NEST (test_gap_junctions_agree_with_nests_waveform_relaxation); a junction between two populations is the same junction within oneNEST iterates waveform relaxation to a tolerance with cubic interpolation of the partners’ voltages; sparx takes one iteration, predicting the partners’ end voltages and interpolating linearly, which is second order and costs a second advance of each coupled population. NEST without waveform relaxation holds the partner’s voltage of the step before and is 0.24 and 2.1 mV off. At 4,000 nS (ten times C / dt) sparx stays bounded, but the fast mode, shorter than a step, is several mV wrong in its transient. A partner that fires is interpolated toward its reset voltage, as NEST interpolates it.
Neuromodulator (Modulator)Its equation, dc/dt = -c / tau + release sum_k delta(t - t_k)The concentration equals the sum of exponentials over the spikes of the named source neurons to 2.3e-7 in float32 at 0.5 ms steps, and a Plasticity rule reads that value after each step (test_a_modulator_follows_its_equation_and_plasticity_reads_it)One concentration per modulator for the whole network, with no spatial extent or receptor kinetics, released by spikes only. DopamineSTDP reads it.
RuleReferenceCheckedDepartures and choices
e-prop (sparx.learn.eprop)Bellec et al. 2020; their numerical verifications (numerical_verification_eprop_*.py in IGITUGraz/eligibility_propagation)Their two identities to 1e-9 relative (tests/test_learn.py): the online gradients equal BPTT’s with the gradient stopped at the recurrent spikes (their stop_z_gradients=True, RecurrentCell(cut_gradient=True) here), and the eligibility traces weighted by the true learning signal equal BPTT’s gradient (their equation 1). Checked for LIF and ALIF with detached resets (ALIF with and without their refractory period of 2 steps, as their verifications run it), LIF with the reset in the gradient, LIF with a decay per neuron, and the current-based Serial(LICell, LIFCell); and at dt = 2 against the same network at dt = 1Their code is TensorFlow 1 and does not run here, so the identities, which they use to verify their own code, are the reference. The eligibility traces are derived for any elementwise model by forward-mode derivatives of its step, not hand-coded per model; a leaky readout’s learning signal is handled by filtering the traces with its leak, which keeps the rule online and exact. A state variable whose input enters with constant coefficients shared by all neurons (the membrane under a detached reset) keeps their filtered presynaptic trace z_bar in place of a vector per synapse, and one no gradient reaches (the refractory count) keeps nothing; which is which is read from the jaxpr of the step’s derivative, and an unclear case keeps the full vector. On SHD’s shapes this is 14 times faster than the vector per state variable for ALIF (benchmarks/bench_eprop.py, performance).
OTTT (sparx.learn.ottt)Xiao et al. 2022; their OnlineLIFNode and WrapedSNNOp (pkuxmq/OTTT-SNN)Every weight and bias gradient of a two-hidden-layer network over 12 steps, against their modules run in PyTorch, to 1e-10 relative (tools/make_ottt_fixtures.py)None found. Their trace starts at the first spike undecayed, which the recursion from zero reproduces. ottt takes the trace’s time constant tau and decays it by exp(-dt / tau); their trace decays by 1 - 1 / tau, which the test passes as the time constant -1 / log(1 - 1 / tau).
REINFORCE (sparx.learn.reinforce, policy_gradient)Williams 1992: grad E[R] = E[R grad log P(s)], for a recurrent layer of BernoulliCellsOn a layer small enough to enumerate (two neurons with recurrence, three steps, 64 trajectories, both reset rules, float64): every eligibility equals the gradient of its trajectory’s log-probability written independently (observed 8.9e-16), the probabilities sum to 1, the expected eligibility is 0, and the probability-weighted sum of reward times eligibility is the exact gradient of the expected reward (observed 1.7e-16). Sampled: policy_gradient over 20,000 trajectories lies within 5 standard errors of that gradient, with and without a baseline at the expected reward, which cuts its spread 3.9-fold (the test asks for twofold). A hundred steps of ascent teach a layer to fire one neuron and silence another. Dropping the zero reset’s derivative fails the enumeration (tests/test_learn.py)The derivative of log P(s_t) in the membrane is beta (s_t - p_t), written from the cell’s sigmoid, not differentiated through it: the ratio (s - p) / (p (1 - p)) is 0 / 0 in float32 once p rounds to 1. The weights’ eligibilities only; the decay, threshold and beta get none.
Reward diffusion and its network (sparx.learn.RNeuralNet, reward_diffusion)RNeuralNet-Research d4b7803: Global_Refresher, Global_ForwardProcessor, Global_Adjuster, Global_RewardSpreader, Global_Teacher and Global_Renew (Processor.cpp), Soma_t, Neuron_t and Neurite_t (neuron.h, neuron.cpp), NeuralNet_init (Main.cpp); the rule as the owner’s research notes reconstruct itThe original’s classes and functions compiled unchanged (g++ 13.3, -O2) with a driver that runs them single-threaded in the order RNeuralNet describes (tools/make_rneuralnet_fixtures.py, tools/rneuralnet_driver.cpp), on 40 neurons of 4 connections each with delays of 1 to 21 ticks, 3 input and 2 output neurons, for 80 ticks with three rewards: every neuron’s output each tick to 1e-6 (observed 7.2e-7 on outputs up to 9.4, on both sides of the thresholds), the local rewards to 1e-7 (observed 6.0e-8) and the weights after each reward to 1e-7 (observed 0). With each delay at Myelin + 1, without the tick a connection loses when its source sends after its target in the pass, outputs differ by 10. The iterative first-visit spread equals a recursive transcription of Global_RewardSpreader on random cyclic graphs (observed 1.8e-7), and depends on the order of the connections as the original does; every path’s spread equals (I - discount P)^-1 r (observed 4.8e-7)The original’s threads share the neurons without synchronization, so RNeuralNet fixes one order of its functions, and calls Global_Adjuster, which Main.cpp leaves commented out. NeuralNet_init makes the output neurons’ connections to the feeder last, which in that order splits their sending in two: wire refuses it and random makes the feeder connection with the neuron’s others. random wires every neuron, where an off-by-one leaves the original’s last one unconnected, and draws from numpy, not rand(). The shares are taken less each target’s largest activity, finite where the original’s float exponentials overflow. The input neurons’ local rewards, which Global_Renew never clears and nothing reads, are not compared.
Fast weights (sparx.dynamics.FastWeights with DecayingHebb, OjaHebb, ModulatedHebb or RetroactiveHebb, on a RecurrentCell; sparx.nn.Recurrent(rule=...))Miconi et al. 2018 (differentiable plasticity, eqs. 1 to 3) and 2019 (Backpropamine, eqs. 3 to 5); their networks: NETWORK of simple/simple.py and Network of maze/maze.py with rule="oja" (uber-research/differentiable-plasticity 5bd29a18), Network of simplemaze/maze.py and of maze/batch.py with type="modul", addpw=3 (uber-research/backpropamine 180c9101)Each of the four against their class run in PyTorch in float64 over 20 to 23 steps (tools/make_miconi_fixtures.py): every step’s activity and trace to 1e-12 (observed at most 1.3e-14 and 4.5e-14, both in the modulated network), the loss to 1e-12 relative, and every parameter’s gradient to 1e-12 of its largest entry (observed at most 1.0e-14). The clip holds part of each bounded trace, and a mutation of each rule fails its case. On their pattern completion, shrunk to two patterns of 8 bits, every rule trained 300 steps gets 4 to 5 % of the zeroed bits wrong, the same network without a trace 22 %; on their full task (1000 bits, five patterns, 2000 episodes; examples/pattern_completion.py) the decaying trace leaves 0.3 % of the zeroed bits wrong and the network without one 50.1 %. A Sparse wiring that lists every pair computes the Dense layer’s activity and traces under every rule (observed 6.3e-7 in float32), and reading a connection’s presynaptic unit for its postsynaptic one fails itA step of dt keeps (1 - eta) ** dt of a decaying trace or eligibility and adds dt times the rate to the other traces; at dt = 1 both are theirs. maze/batch.py keeps its matrices [post, pre], transposed here. Their tanh unit is RateCell(0.0); the Rate layer also learns a bias per unit, which their networks take from a bias neuron or an input layer. The fan-out of ModulatedHebb is in their code and their appendix, and RetroactiveHebb has none, as maze/batch.py has none. sparx.nn’s neuromodulator starts as their torch.nn.Linear layers do
Predictive coding and PC-ALM (sparx.learn.PredictiveCoding, residual_mlp, PredictiveCodingObjective)Seely and Gould 2026 (arXiv 2605.31022), eqs. 8, 11 and 12 and Algorithm 1; their JAX reference, pcalm.inference and pcalm.model (SakanaAI/pc-alm 660747f6)Against their code in float64 (tools/make_pcalm_fixtures.py) on two of their residual MLPs (tanh, depth 4; ReLU, depth 8 at T = 2L), for PC and for PC-ALM with either weight-credit timing and with two activity steps per multiplier step: every settled activity to 1e-12 (observed 8.9e-16), every multiplier to 1e-12 (observed 1.6e-15), and every weight update to 1e-12 of its layer’s largest entry (observed 4.8e-14). A mutation of the multiplier’s sign, of the credit timing, of the number of cycles, of the state the update reads, or of the per-example step fails its cases. Backpropagation through residual_mlp gives their BP gradient, and the objective hands the rule’s update to dew’s trainer unchanged (observed 0). Their headline cell (Fashion-MNIST, width and depth 32, one epoch) trained through dew at seeds 0, 1 and 2 scores 75.1, 76.1 and 76.5% by PC-ALM, 62.2, 65.5 and 65.6% by PC and 77.8, 76.9 and 76.7% by BP, against their code on the same machine at 77.73, 76.49 and 76.34%, 68.17, 64.49 and 66.54%, and 78.65, 76.85 and 77.31%Each example descends its own energy at state_lr, the step their code reaches by scaling the batch mean’s gradient by the batch size; the weight update is the energy summed over the batch’s real rows, which dew’s trainer divides by their count, their mean. The weights start from dew’s draw of their distribution, so a run matches theirs in distribution, not draw for draw
Event-based exact gradients (sparx.learn.events)Wunderlich and Pehle 2021 (EventProp): the exact gradient of LIF networks with exponential current synapses in continuous timeGradients by weights and by input spike times against central differences of the exact simulation, to 1e-6 relative in float64 (test_event_gradients_are_the_exact_derivatives_of_spike_times); the event simulation fires as sparx.dynamics’ LIF integrated at 1 us does, spike for spike within 3 us, and a burst of 19 spikes between two inputs fires every spike the grid does, each a threshold crossing of the closed-form membrane to 1e-12 (test_a_strong_pulse_fires_every_spike_until_capacity_runs_out); a two-layer network learns a latency taskThe gradient is EventProp’s, obtained differently: each threshold crossing is found by bisection on the membrane’s monotone rising branch and then differentiated by the implicit function theorem (one Newton step with the root held fixed), and autodiff of the event loop does the rest, where EventProp integrates adjoint dynamics backward between the spikes. Both are the exact derivative with spike counts fixed; this one stores each event of the forward pass, EventProp’s adjoint stores only the spike times. Their reference code was not run.
ANN-to-SNN conversion (sparx.learn.convert)Rueckauer et al. 2017 (robust normalization, IF with reset by subtraction, analog input, batch norm folding, average pooling, max pooling gated by the most active input); their toolbox, snntoolbox 0.6Their rate bound: a neuron under a constant drive a in threshold units fires at clip(a, 0, 1) within 1 / T (T = 10, 100, 1000). A ReLU MLP converted with the 99.9th percentile reaches its accuracy within 2% at 300 steps. A CNN with batch norm, max and average pooling keeps the ReLU network’s predictions on all 500 test images at 300 steps, and hidden rates correlate with normalized activations above 0.99. Folding batch norm reproduces flax’s outputs to float precision. Against snntoolbox run on a Keras CNN (tools/make_snntoolbox_fixtures.py): the folded and normalized weights match to float32 precision, the first layer fires the same spikes, and predictions agree on all 50 test images (output rates correlated at 0.99991). A converted MLP with reset to zero exports to NIR and imports back with the same weights and spikesAverage pooling passes averaged spikes to the next layer; snntoolbox puts IF neurons on it, which differs by up to 1 / T. The max-pool gate counts spikes before the current step; snntoolbox counts the current step too, which lets a pool outfire its most active input (by up to 0.21 on the fixture), so spikes after the first max pool differ (with their rule, the pool and the next layer fire the same spike counts as theirs). Unpadded pooling and 2D NHWC convolutions only; the spiking softmax is not converted. The converted network flattens with sparx.nn.Flatten, channels first; the toolbox’s dense layer reads Keras’s channels-last flatten of a 1 x 1 map, where the two orders agree, and the test reorders its rows explicitly all the same.
FormatReferenceCheckedDepartures and choices
NIR (sparx.nir)NIR 1.0.8; snnTorch 1.0.0’s export_to_nir and import_nirThree networks exported by snnTorch (two dense layers; two Conv2d layers with Flatten; RLeaky) import into sparx and fire spike for spike as snnTorch ran them, hidden layers included; exported back, every weight, bias, time constant, threshold, stride, padding, dilation and edge equals snnTorch’s graph; a sparx stack of convolutions, Flatten, Recurrent(LIF) and dense layers round-trips through NIR to identical spikes, and a classifier with an LI readout to identical membranes under both discretizations (tests/test_nir.py)NIR’s neurons are continuous, and each library picks its step: snnTorch reads them by forward Euler at a fixed 1e-4 s. Sparx states the discretization on both import and export ("exact", its own, or "euler"). NIR is channels first and sparx channels last; weights are transposed and sparx.nn.Flatten orders features C, H, W as PyTorch does. snnTorch’s RLeaky feedback has a bias that sparx’s lacks; it moves into the previous layer’s bias and exports back as a Linear node. snnTorch’s import_nir fails on recurrent graphs, so sparx’s recurrent export is checked against its own import. NIR 1.0.8’s type check rejects non-square and grouped kernels; sparx imports them from an unchecked graph only. snnTorch’s exporter writes scalar parameters with a shape NIR’s type check rejects; the fixtures give per-neuron parameters. Only hard-reset LIF and IF neurons, and LI readouts, so far; an IF node’s r is 1 / dt under either discretization. NIR’s LI is its LIF without a threshold, so it shares the LIF’s discretization, which the snnTorch fixtures check; no library’s own LI export is compared.
FunctionReferenceCheckedDepartures and choices
encode.LatencyEncodersnnTorch spikegen.latency(linear=True, normalize=True, clip=True)ExactsnnTorch fires values below the threshold at the last step unless clip=True; sparx never fires them.
encode.DeltaEncodersnnTorch spikegen.delta(padding=False)Exact, with and without off-spikesNone.
losses.per_step_cross_entropysnnTorch ce_rate_lossWithin 1e-6 relativeNone.
losses.rate_msesnnTorch mse_count_lossWithin 1e-6 relative, as snnTorch’s value divided by TsnnTorch’s squared error of counts over T grows with T; sparx’s rates do not, and snnTorch rounds its target counts down.
losses.van_rossumvan Rossum 2001; Elephant 1.2.1 van_rossum_distanceAll pairs of six random trains at three time constants, to 1e-10 (tools/make_elephant_fixtures.py)Elephant’s distance is sqrt(2) times van Rossum’s: it drops the 1/2 of the closed form. Sparx keeps the paper’s definition. The integral is exact for spikes on the grid, the tail after the last step included, not a Riemann sum.
spiketrains.victor_purpuraVictor and Purpura 1996; Elephant victor_purpura_distanceAll pairs, three costs, exactlyNone.
spiketrains.coincidence_factorKistler et al. 1997, in Jolivet and Gerstner’s (2004) form; brian2modelfitting 0.4 get_gamma_factor(rate_correction=False)21 pairs of trains over 2 s at three precisions: the same coincidences, within 4e-15, and the same Γ where the spike counts are equal, within 6e-17 (tools/make_gamma_fixtures.py)brian2modelfitting takes the chance rate ν from the reference train. Jolivet and Gerstner take the predicting train’s, under which a Poisson train of any rate scores 0 on average, and sparx follows them. A neuron silent in both trains has no Γ (NaN).
SurrogatesThe papers’ formulasGradients and forward-mode tangents, and each surrogate’s areaNone.
losses.softmax_sum_cross_entropy, readout="softmax_sum"SNN-delays’ calc_loss with loss='sum' and its predictionTheir formula in float64, within 1e-6 relative, and through their whole network (above)None: their summed probabilities go through torch’s CrossEntropyLoss, which takes a log-softmax of them again, and sparx does the same.
datasets.bin_events(binning="events"), shd(binning="events")SpikingJelly’s integrate_events_by_fixed_duration_shd in the 2023 tree SNN-delays runs onExact on a recording with silences (test_event_binning_reproduces_snn_delays_frames), and on all 10,420 SHD recordings against the frames their running code cached, at 10 ms with channels summed in fivesTheir frames open at events: each takes the events within 10 ms of the first one not yet counted, so silences longer than a step vanish and recordings shorten (at most 124 steps where the grid needs 137). Later SpikingJelly releases bin on a grid from the first spike. Times are scaled to ms in the file’s float16, as theirs are.
PieceReferenceCheckedDepartures and choices
dew’s OneCycle, Cosine with no warmup and Exponential, stepped once an epoch (every)SNN-delays’ schedulers: torch’s OneCycleLR(max_lr=5e-3) for the weights and its Adam momentum cycle, CosineAnnealingLR for the positions, decrease_sig for the width, each stepped at the end of an epochAll 150 epochs of their SHD configuration, recorded from their code: learning rates within 1e-5 relative (the cosine within 1e-4), momentum within 1e-6, width within 1e-5 (test_dews_schedules_stepped_once_an_epoch_are_snn_delays_torch_schedulers)None.
dew’s OptimConfig(optimizer="adam") with a ParamGroup each for SNN-delays’ delays, weights and normstorch’s Adam(weight_decay=...), whose decay is an L2 term added to the gradient, with OneCycleLR setting its lr and betas[0] before each step, and DCLS’s clamp of the positionsDew’s own tests, against torch’s update (tests/test_optim_schedules.py in dew); sparx checks that the groups find a DelayedDense’s delays and keep them within their bounds (test_dews_parameter_groups_find_the_delays_and_keep_them_within_bounds)Groups are matched by path patterns; torch takes lists of tensors.
examples/train_shd.py --recipe snn-delaysTheir best_config_SHD.py and training loopThe network (214,016 parameters, their count), loss, schedules and binning, as above. The first 2 epochs of the 20-epoch schedule on SHD with a 10% holdout: test accuracy 8.7% then 19.8% (validation 9.4%, 17.9%), where their code on the same schedule and all of the training set scores 8.1% then 23.5%; 7.5 to 14 s a training step of 256 on a shared 4-core CPUTheir batches are padded to the batch’s longest recording (about 105 steps on average); sparx pads every recording to 124, so the summed softmax and the batch statistics also see the extra silent steps. dew drops an epoch’s last partial batch (31 steps of 256, where they take 32 with a last batch of 220). Their test accuracy is the mean of per-batch accuracies; sparx’s is over all 2264 recordings. They select epochs on the test set and report the best test accuracy; the example holds out 10% of the training set by default and reports the test accuracy at the best validation epoch beside theirs. Random streams (initialization, shuffling, dropout) differ.
  • Shiu et al.’s model on the male CNS is calibrated, not validated: a published simulation on this connectome, or more behaviours than sugar to MN9, would test it. The left and right sugar neurons drive MN9 unequally.