orbitsight / docs /OrbitSight_Full_Technical_Proposal_Review.md
HishaamA's picture
Add DAVIS animation to review gallery
53aa020 verified
|
Raw History Blame Contribute Delete
53.5 kB

OrbitSight — Full Technical Proposal and Engineering Review

Long-form review draft, 29 August 2026
Team SIMER — Hishaam A., Simreen Siraj, and Christian Sabatini — Ampiere Labs

This is deliberately not the five-page submission proposal. It is the detailed source document from which a shorter competition proposal, pitch deck, or technical appendix can be produced. It tells the complete development story, including failed models, invalid experiments, latency failures, and the limits of the available evaluation evidence.

Executive summary

OrbitSight detects and tracks resident space objects (RSOs) directly from asynchronous neuromorphic vision sensor events. The final submission model is a compact recurrent neural network followed by a physically interpretable Kalman tracker. It reads raw *.npy event recordings, processes consecutive 40 ms windows, and writes the exact ChallengeON prediction files and evaluation workbook from an automatic, offline, CPU-only Docker container.

The important story is not a single final score. It is a sequence of engineering discoveries:

  1. The first honest pipeline scored only F1 0.036 / mAP 0.010. Early capped experiments had silently missed most labelled windows, and several apparently good measurements were not reproducible.
  2. Correct windowing, event denoising, a three-channel representation, recurrent detection, CIoU localization, longer temporal training, and honest best-epoch selection produced V5. After output calibration and gate corrections, V5 reached approximately F1 0.630 / mAP 0.380 on the leakage-free local diagnostic.
  3. V7 introduced a more principled area-resize and coordinate-mapping contract. It did not beat V5 overall, but it was substantially better on EVK4. V5 remained better on DVX.
  4. A sensor router combined the specialists: EVK4 → V7, DAVIS/DVX → V5. This hybrid reached F1 0.668 / mAP 0.490 leakage-free and 0.676 / 0.549 across all four local recordings. The route was chosen after per-sensor results were visible, so it is a score-maximizing diagnostic rather than blind model-selection evidence.
  5. The heavy hybrid was unusable on the competition CPU: approximately 383 ms mean / 451 ms p95, around an order of magnitude above the 40 ms requirement.
  6. The hybrid therefore became a teacher rather than a deployable model. A 135,111-parameter recurrent student was trained from human labels plus carefully filtered teacher evidence. It retained full 640×640 localization while being approximately 63.6× smaller than one 8,588,339-parameter teacher branch.
  7. The distilled student measured F1 0.6229 / mAP 0.4608 leakage-free and 0.6375 / 0.5373 on all four recordings, with 18.55 ms mean / 21.97 ms worst-sequence p95 in a full local CPU profile.
  8. Without retraining, conservative sensor calibration and tracker changes raised the final diagnostics to F1 0.6292 / mAP 0.4645 leakage-free and 0.6437 / 0.5410 all-four. A later 3,920-window profile measured 17.05 ms mean / 20.64 ms worst-sequence p95, with the largest observed window at 39.17 ms.

The result is not the highest-scoring network we built. It is the strongest practical balance of detection quality, CPU latency, reproducibility, and challenge-format correctness.

1. Problem and evaluation objective

1.1 What the sensor produces

A conventional camera exposes a dense image at fixed intervals. A neuromorphic vision sensor works differently: each pixel emits an event when local log-brightness changes enough. An event carries pixel position, polarity, and a microsecond timestamp. The result is a sparse, asynchronous stream rather than a sequence of ordinary photographs.

This is attractive for space situational awareness because event sensors offer high dynamic range, low motion blur, and fine timing. It is also difficult because most events are not the target. In the supplied recordings, labelled RSO events represent roughly 0.08–0.96% of the stream. The rest includes sensor activity, hot pixels, star-field responses, and other background structure.

OrbitSight supports three sensor families with very different native resolutions:

Sensor Native resolution Transformation challenge
DAVIS346c 346×260 Must be enlarged while preserving tiny target geometry
DVXplorer 640×480 Near the model scale, but strongly affected by low-SNR clutter
EVK4 1280×720 Must be reduced without deleting sparse events

1.2 Why the targets are unusually difficult

  • The median labelled box is only about 10×11 pixels; the tenth percentile is about 6×6 pixels.
  • IoU ≥ 0.5 on a ten-pixel object allows only a few pixels of center error.
  • Signal strength varies by roughly 200× across recordings.
  • Every annotated window contains one target box, so a second emitted detection cannot become another true positive; it can only reduce precision.
  • The hardest low-SNR cases contain only about six useful object events in a 40 ms window. At that point, localization is limited by sampling noise as much as by model capacity.

1.3 What must be delivered

For every input sequence, the submission produces <sequencename>.txt with these fields:

sequence_id, window_start_timestamp_us, window_end_timestamp_us, x_centre, y_centre, w, h, class_id, confidence

The container also produces Evaluation_Metrics.xlsx. Accuracy is reported using precision, recall, F1, and AP/mAP at IoU 0.5. The competition additionally calls for end-to-end latency below 40 ms on an Intel Core i9-12900H CPU. There is no track-ID column, so tracking matters only when it improves the emitted detection stream.

Literature review and design lineage

OrbitSight is an application-specific synthesis rather than a direct reproduction of one published network. Its design draws on neuromorphic sensing, recurrent event detection, anchor-free dense detection, model compression, and detection-assisted tracking.

Event cameras and space situational awareness

Gallego et al. formalize the event-camera model: a pixel reports a change in log intensity rather than an absolute image sample. The literature emphasizes microsecond timing, high dynamic range, low motion blur, and sparse output, while warning that frame algorithms require an event representation or event-native computation. OrbitSight preserves polarity and within-window timing, but packs events into 40 ms tensors so convolutional and recurrent operations run efficiently on a CPU. Gallego et al., Event-based Vision: A Survey, TPAMI 2020.

Afshar et al. demonstrated event-based RSO detection and tracking across sensors and observing sites. Ralph et al. subsequently studied astrometric calibration and source characterization for event-based space imaging. These works motivate OrbitSight's explicit sensor profiles, native-coordinate reconstruction, and treatment of preprocessing and coordinate mapping as part of the model contract. Afshar et al., Event-based Object Detection and Tracking for Space Situational Awareness, 2019; Ralph et al., Astrometric Calibration and Source Characterisation of Neuromorphic Event-based Cameras for Space Imaging, 2022.

Recurrent event-based object detection

Perot et al.'s Recurrent Event-camera Detector showed that direct event representations combined with recurrent memory outperform feed-forward event detectors because state integrates evidence through time. RVT later showed that multi-stage recurrent backbones can obtain a favorable latency/accuracy tradeoff. OrbitSight adopts direct event tensors, multi-scale features, and persistent temporal state, but uses narrow convolutions and ConvLSTM instead of a transformer for an offline CPU target and tiny objects. Perot et al., Learning to Detect Objects with a 1 Megapixel Event Camera, NeurIPS 2020; Gehrig and Scaramuzza, Recurrent Vision Transformers for Object Detection With Event Cameras, CVPR 2023.

ConvLSTM replaces fully connected state transitions with spatial convolutions, retaining where weak evidence occurred. OrbitSight places one after every downsampling stage and uses dilations 1, 2, 4, and 8 to enlarge spatial context without changing tensor shapes. Shi et al., Convolutional LSTM Network, NeurIPS 2015.

Anchor-free detection, imbalance, and localization

FCOS established fully convolutional anchor-free prediction; YOLOX paired anchor-free boxes with a decoupled head. OrbitSight adopts those ideas with custom assignment and scoring for one RSO class. It predicts raw center offsets and log sizes on three scales without an anchor bank. Tian et al., FCOS, ICCV 2019; Ge et al., YOLOX, 2021.

Dense event tensors create tens of thousands of background cells for at most one target. Focal loss suppresses easy negatives so hard examples control learning. OrbitSight uses sigmoid focal BCE for objectness. CIoU combines overlap, center distance, and aspect-ratio consistency, aligning training more closely with the IoU≥0.5 metric than independent coordinate L1. Lin et al., Focal Loss for Dense Object Detection, ICCV 2017; Zheng et al., Distance-IoU Loss, AAAI 2020.

Distillation and tracking

Knowledge distillation transfers useful behavior from an expensive model or ensemble into a smaller deployable network. OrbitSight follows that deployment logic but does not blindly copy teacher output: human boxes remain authoritative, matched teacher boxes provide low-weight targets, and proposals in human-empty windows become negative evidence. Hinton, Vinyals, and Dean, Distilling the Knowledge in a Neural Network, 2015.

ByteTrack showed the value of associating lower-confidence detections. OrbitSight evaluated that two-tier idea, but production emits only stable global top-one at 0.10, making the nominal low tier unreachable. The final system is therefore a single-tier constant-velocity Kalman/coast tracker using the same association machinery. Zhang et al., ByteTrack, ECCV 2022.

OrbitSight's design contribution

  • a polarity/recency representation with sensor-specific event denoising;
  • a full-resolution but ultra-narrow recurrent detector;
  • single-class objectness decoding without a redundant class multiplier;
  • selective, human-authoritative distillation from a sensor-routed teacher;
  • weights bound to preprocessing and coordinate-mapping fingerprints;
  • exact stable top-one decoding that avoids NMS without changing the selected box;
  • one evaluated detector/tracker/writer deployment contract.

2. Development history at a glance

The following table separates comparable leakage-free results from historical development measurements. The detector_e30 row came from the validation split used at that time and is therefore useful as progress evidence, not a strict comparison with later test rows.

Stage Main change Leakage-free F1 Leakage-free mAP Decision
First honest baseline Correctly measured initial pipeline 0.036 0.010 Foundation only
detector_e30 30-epoch trained pipeline and tuned tracking 0.515* 0.220* Promising, not directly comparable
detector_ciou CIoU, dilations, regime balancing 0.5620 0.2536 Retained ideas
V5 detector TBPTT-40, honest split, best epoch 0.5808 0.2749 Strong detector checkpoint
V5 + box calibration Per-sensor size correction 0.5933 0.2947 Retained
V5 complete pipeline Low detector gate and per-sensor emit gates ≈0.6300 ≈0.3798 Scientific baseline
V6 Intended area-resize successor 0.6116† 0.3570† Quarantined: invalid contract
V7 epoch 33 Correct area resize and pixel-center mapping 0.5954 0.3585 Lost overall; EVK4 specialist
Routed V5/V7 hybrid EVK4 V7, DAVIS/DVX V5 0.6683 0.4895 Accuracy teacher; too slow
Compact student from scratch Same small architecture, no teacher 0.1055 aggregate 0.2602 aggregate DVX collapsed
Distilled CPU student Human labels + selective hybrid teaching 0.6229 0.4608 Fast deployable candidate
Final CPU student No-retrain calibration and tracker refinement 0.6292 0.4645 Recommended submission

* Historical validation regime.
† Retained only as an audit trail because training inputs and declared preprocessing did not match.

Measured model evolution

Two lessons are visible immediately. First, the large gain from below 0.1 to approximately 0.63 did not come from one architectural trick; it came from correcting the entire data, training, inference, tracking, and writing contract. Second, the best accuracy model was not the best deployment model. The heavy hybrid forced a separate compression phase.

3. Why the first scores were below 0.1

3.1 The first honest result

The first reproducible leakage-free baseline was F1 0.0364 / mAP 0.0096. This was useful because it established an honest floor, even though the score was extremely low. An earlier F1 value around 0.14 came from a measurement path that could not be reproduced consistently and is not used as the project baseline.

3.2 The label-starvation bug

Early capped experiments evaluated the first N windows of a stream. Labels, however, often began much later. In one sequence the first labelled window was 1,562, so a 1,500-window cap could finish before reaching the main annotated region. Across the affected capped runs, only about 34% of labels were visible.

This bug mattered twice: training batches contained too few positive examples, and evaluation looked at the wrong temporal region. The fix introduced an explicit starting window and aligned capped runs to the first labelled window. This was not a model improvement; it was a measurement correction that made later model improvements meaningful.

3.3 Other foundational corrections

  • Ground-truth timestamps were required to lie on the exact 40 ms window phase.
  • The official evaluator's pixel-inclusive IoU convention was reproduced locally.
  • Prediction box fields were made integer-valued only at the final writer boundary.
  • Zero-sized boxes after rounding were clamped to at least one pixel.
  • Model objectness was decoded correctly for the single-class RSO task instead of being multiplied by a redundant class probability.
  • Detector, recurrent, and tracker state were reset at every sequence boundary.

These corrections are easy to underestimate. A sophisticated model evaluated through a misaligned or inconsistent pipeline is still a bad system.

4. The common event-processing pipeline

The heavy models and compact student share the same conceptual flow:

raw event stream
    ↓
memory-mapped ingestion and phase-checked 40 ms windows
    ↓
sensor-specific background-activity filtering
    ↓
three-channel event tensor at 640×640
    ↓
recurrent anchor-free detector
    ↓
map box back to native sensor coordinates
    ↓
Kalman association and bounded coasting
    ↓
sensor-specific output gate and one-box cap
    ↓
integer portal row + Evaluation_Metrics.xlsx

4.1 Background-activity filtering

A real optical signal tends to produce neighboring events close together in time. Independent sensor noise is less likely to have that support. The BA filter asks a simple question for each event: did enough nearby pixels fire recently? The parameters are sensor-specific:

Sensor Required support Radius Time horizon RSO events retained Noise rejected
DAVIS 1 6 px 40 ms 99.47% 28.32%
DVX 1 4 px 20 ms 98.24% 39.94%
EVK4 1 3 px 5 ms 97.71% 41.41%

The filter is window-local and stateless in the deployed route. Hot-pixel support exists but is not silently enabled because changing preprocessing only at inference would create a train/deployment mismatch.

4.2 Three-channel representation

Each 40 ms event window becomes a three-channel tensor:

  1. positive-event count;
  2. negative-event count;
  3. recency of the latest event.

The first two channels show where brightness increased or decreased. The recency channel preserves temporal direction inside the window. A moving point source often produces a polarity pattern and timestamp gradient, so these three channels retain more useful physics than a single accumulated image.

The model input remains 640×640 even in the compact student. Reducing the image would save compute, but a 6–11 pixel target cannot tolerate much spatial quantization. We reduced network width instead of discarding localization resolution.

5. From early trained models to V5

5.1 The detector family

OrbitSight uses an anchor-free, multi-scale detector. “Anchor-free” means that the network predicts a target center and size directly at grid locations instead of choosing from a bank of predefined rectangle shapes. This suits small RSOs whose apparent extent depends on brightness, motion, optics, and sensor resolution.

The backbone contains four downsampling stages. Each stage includes ConvLSTM recurrence, so the model carries information from earlier windows. Recurrent dilations (1, 2, 4, 8) expand the spatial context without requiring a much deeper network. Detection heads operate at multiple scales, with stride 4 retained for precise small-target localization.

5.2 CIoU and confidence learning

Early box regression did not optimize the competition's overlap criterion closely enough. Complete IoU (CIoU) improved the learning signal by combining overlap, center distance, and shape consistency. It also supplies a gradient when a predicted box and target do not yet overlap.

The objectness target is linked to achieved localization quality. In plain language: a box should be confident when it is both likely to contain the RSO and accurately localized. This makes confidence useful for ranking, but it also means correct tiny boxes may have moderate rather than near-one scores. That observation later motivated separating the detector's low proposal threshold from the final output gate.

5.3 Training improvements that produced V5

  • Regime balancing: batches were designed to prevent abundant background windows from erasing rare positive examples.
  • TBPTT length 40: the network learned across 1.6 seconds of 40 ms windows instead of a shorter temporal horizon.
  • Honest validation: sequence-level separation reduced leakage from adjacent windows.
  • Best-epoch selection: V5 kept epoch 33 rather than simply shipping the final epoch.
  • Full-path validation: evaluation included tracking, coasting, gates, and writer quantization—the path that would actually be submitted.

V5's heavy backbone widths are (32, 64, 128, 256), the head width is 128, and the model has 8,588,339 parameters. It became the scientific baseline because it was the strongest model supported by a coherent, reproducible contract.

5.4 The output-gate breakthrough

Originally, the detector threshold and final output threshold were effectively the same. A box below the emit threshold was discarded before tracking, even if it was the best proposal in the window. The corrected design runs the detector at a low 0.10 proposal threshold, lets the tracker use that evidence, and applies the sensor output gate only at the writer.

This separation was one of the largest no-retrain gains in the heavy pipeline. The retained V5 gates were DAVIS 0.30, DVX 0.10, and EVK4 0.10. Per-sensor box-size correction also repaired systematic under-sizing caused by sensor transformations. Together, these changes moved the complete V5 pipeline to approximately F1 0.630 / mAP 0.380 leakage-free.

6. V6 and the importance of a preprocessing contract

V6 was intended to test area resampling and add dim-regime data. Its recorded score appeared competitive, but an audit found a critical mismatch: training consumed historical nearest- resized tensors while metadata claimed area resize. Evaluation could therefore feed one representation to weights trained on another.

The checkpoint still loaded because tensor shapes matched. That is exactly why the problem was dangerous: shape compatibility did not imply semantic compatibility. V6 was quarantined, and later loaders were strengthened to bind checkpoints to preprocessing fingerprints.

The general lesson is central to OrbitSight: a model is not just weights. It is weights plus window definition, denoising, resize method, coordinate mapping, count transform, state behavior, thresholds, and writer projection.

7. V7: better geometry, mixed sensor behavior

V7 rebuilt the preprocessing contract around true area resampling and pixel-center rounding. The intent was sound: when EVK4 is reduced from 1280×720, nearest-neighbor sampling can delete sparse events, while area resampling preserves aggregate activity more faithfully.

The consolidated V7 model did not beat V5 overall:

Model Leakage-free F1 Leakage-free mAP All-four F1 All-four mAP
Paired V5 0.6298 0.3804 0.6403 0.4670
V7 epoch 33 0.5954 0.3585 0.6100 0.4378

However, aggregate results hid meaningful specialization:

Sequence V5 F1 / AP V7 F1 / AP Better branch
EVK4 magnitude 7.3 0.5175 / 0.2797 0.7121 / 0.6070 V7 by a large margin
DVX Stars3 0.6667 / 0.5812 0.5723 / 0.3525 V5
DVX Thuraya3 0.4357 / 0.2803 0.2095 / 0.1160 V5

V7 was therefore not a universal successor. It was an EVK4 specialist. This distinction created the hybrid teacher.

8. The V5/V7 sensor-routed hybrid

The router chooses one heavy branch once per sequence:

  • EVK4 → V7 epoch 33;
  • DAVIS → V5 best;
  • DVX → V5 best.

Only one branch runs for a given window, so this is routing rather than simultaneous ensembling. The routing result was:

Scope Precision Recall F1 mAP
Leakage-free 0.6400 0.6991 0.6683 0.4895
All four 0.6504 0.7036 0.6760 0.5489

These were the highest practical local accuracy values in the project. They are also post-hoc: the route was proposed after the per-sensor results were known. The hybrid is therefore useful as a score-maximizing package and as a teacher, but its local result is not an unbiased estimate of unseen performance.

8.1 Why the hybrid could not be submitted on CPU

Each branch still has 8.59 million parameters and full-width recurrent computation. On a local four-thread CPU, the hybrid measured approximately:

  • 383 ms mean per window;
  • 451 ms p95 per window;
  • roughly 313 ms/window effective during the complete four-sequence replay.

The budget is 40 ms. This is not a small optimization gap. Tracker tuning, output formatting, or minor Python changes cannot create a tenfold speedup in the dominant convolutional model. The architecture had to become much smaller.

9. Why teacher–student distillation was chosen

The hybrid knew useful sensor-specific behavior, but was too slow to deploy. Distillation separates those roles:

  • the teacher is allowed to be large and expensive because it runs only while generating training evidence;
  • the student learns from human labels plus selected teacher information, then runs alone in the submission container.

The shipped image does not contain V5, V7, or the router. Teacher latency is therefore absent from deployment latency.

9.1 Why a small model trained from scratch was not enough

The same compact architecture was first trained without the routed teacher. Its aggregate F1/mAP was 0.1055 / 0.2602 and DVX effectively collapsed. On a six-positive DVX held-out capture, the best examined epoch produced 2 TP, 3,941 FP, and 4 FN. The two true positives ranked 793rd and 2,595th by confidence.

No output threshold can repair that ordering. Raising the threshold removes the buried true positives along with false positives; lowering it emits thousands of false detections. The failure was a hard-negative ranking problem, which is exactly where a stronger teacher can provide useful structure.

9.2 Selective distillation rather than blind imitation

Teacher outputs are not automatically true. The training design keeps human ground truth authoritative:

  • A teacher detection that matches human GT may contribute a low-weight soft distillation target.
  • A human GT missed by the teacher remains a full positive.
  • A teacher detection in a human-empty window is treated as hard-negative/ranking evidence, not copied as a positive.
  • Post-tracker teacher boxes are diagnostics only; the student learns pre-tracker detection behavior so it is not tracked twice.
  • Teacher recurrence, student recurrence, BA state, and tracking state reset at sequence boundaries.
  • Cache manifests bind source hashes, sensor parameters, preprocessing, coordinate mapping, checkpoint identity, and code/schema versions.

This design prevents the most dangerous distillation failure: teaching the student to repeat the teacher's false positives.

9.3 Training evidence and run

Teacher evidence was generated in two independent shards and merged under signed manifests. Across the merged cache there were:

  • 17 training sequences;
  • 106,194 windows;
  • 18,139 pre-tracker teacher detections;
  • 16,682 post-tracker diagnostic detections;
  • 11,343 positive windows with a matching teacher detection;
  • 5,238 human-empty windows containing teacher proposals, retained as hard negatives.

The final A100 training used 36 epochs, 260 updates per epoch, 9,360 optimizer updates, batch 16, and segment length 20. An interrupted client connection was recovered by exact checkpoint resume, including model, optimizer, scheduler, RNG, sampler, and batching state. The final checkpoint SHA-256 is:

5f176b9baffc3478b1aaada13277e2c6093ba02439b76b8330cb8d654f5f554d

10. Final-model methodology

This section describes how the final student was produced and evaluated. Section 11 then defines the frozen network and inference path component by component.

10.1 Research questions and temporal unit

Development asked whether one model could localize tiny RSOs across all three sensor families, inherit useful sensor-specific behavior from the heavy hybrid, and keep the complete deployed path below 40 ms. The atomic example is a half-open 40,000 µs event interval. Window phase is anchored to the sequence timestamp contract. Recurrent, BA, and tracker state are initialized at the start of a sequence, advanced exactly once per window, and destroyed at the boundary.

Training uses ordered segments, not shuffled independent frames. Final distillation used 20 consecutive windows per segment—0.8 seconds of temporal context—and batches of 16 independent segments. Each batch member carries its own four hidden/cell state pairs before truncated backpropagation detaches them.

10.2 Data preparation

For every window the method:

  1. memory-maps and timestamp-slices events;
  2. applies the sensor-specific neighborhood/time BA rule;
  3. builds positive-count, negative-count, and latest-event-recency planes;
  4. area-resizes the planes to 640×640;
  5. maps human boxes with the same pixel-center convention;
  6. builds the optional per-event auxiliary target.

Counts remain raw. Stateful BA, sensor FiLM, hot-pixel changes, and target-event thinning are disabled in the final route. Cache fingerprints bind the source, sensor profile, 40 ms phase, input size, count transform, area resize, and coordinate mapping. This prevents a shape-compatible checkpoint from being used with semantically different preprocessing, the failure that invalidated V6.

10.3 Human-authoritative teacher evidence

The V5/V7 router runs offline. Its pre-tracker proposals are stored in sidecars aligned one-to-one with student windows and marked as matched or unmatched against human GT. Post-tracker teacher boxes are diagnostic only. The selective objective is:

L_total = L_human
        + 0.15 L_teacher-objectness-on-matched-GT
        + 0.25 L_teacher-box-on-matched-GT
        + 0.35 L_teacher-empty-window-negative

Teacher scores below 0.05 are ignored. Matched proposals supply a BCE confidence target and Smooth-L1 encoded-box target. A proposal in a human-empty window pushes the corresponding student objectness toward zero. An unmatched teacher proposal never creates a positive label. The teacher is therefore a ranking guide and hard-negative miner, not a replacement annotator.

10.4 Optimization protocol

The final A100 run used 36 epochs, 260 updates per epoch, and 9,360 optimizer steps:

Item Final setting
Optimizer Adam
Initial / minimum LR 5×10⁻⁴ / 1×10⁻⁵
Warm-up / schedule 100 steps / cosine decay
Precision FP32; TF32 disabled
Gradient-norm clip 1.0
Batch / segment 16 / 20 windows
Negative/positive target 2.0
Maximum oversampling 4.0×
Human box loss pixel-inclusive CIoU, weight 2.0
Auxiliary segmentation BCE, weight 0.5
Objectness target floor 1.0 throughout
Quality head / loss disabled / 0.0
Seed 1337

Adam state, scheduler, RNG, batching order, partial epoch progress, preprocessing fingerprint, and the training contract were included in exact-resume checkpoints. A remote disconnect therefore resumed the same stochastic run rather than launching a similar but different experiment.

10.5 Evaluation methodology

Predictions are scored only after native mapping, calibration, tracking, clipping, top-one selection, sensor output gating, and integer writer projection. IoU uses the official pixel-inclusive convention. Metrics are reported for the three-sequence leakage-free aggregate and all four local recordings. Because every remaining recording has influenced diagnostics, these are reproducibility and local comparison results—not blind-generalization proof. Latency is measured after warm-up at streaming batch one and reports mean, per-sequence p95, and maximum rather than forward-pass mean alone.

11. Final model architecture: complete component reference

11.1 Frozen architecture contract

Property Final value
Class OrbitSightDetector
Inference input 1×3×640×640
Backbone widths 4, 8, 16, 32
ConvLSTM dilations 1, 2, 4, 8
Detection scales stride 4, 8, 16
Head width 16
Classes one RSO foreground class
Stored parameters 135,111
Quality head / sensor FiLM absent / absent
Auxiliary segmentation stored; training-only execution
Classification towers stored for compatibility; never executed

Exact final student architecture

11.2 Compression decision

Component Heavy V5/V7 branch CPU student
Input 3×640×640 3×640×640
Backbone widths 32, 64, 128, 256 4, 8, 16, 32
Head width 128 16
Dilations 1, 2, 4, 8 1, 2, 4, 8
Parameters 8,588,339 135,111
Runtime role Offline teacher Deployed model

Heavy teacher versus compact student

The 63.6× reduction comes mainly from channel width. Input resolution, recurrence, the finest stride-4 grid, and three detection scales remain because spatial quantization is expensive for 6–11 pixel targets.

11.3 Input representation

X ∈ R^(B×3×640×640) is float32. Channel 0 counts positive-polarity events; channel 1 counts negative events; channel 2 records normalized recency of the latest event. Empty pixels are zero. This preserves brightness-change sign and within-window order without reconstructing an intensity image.

11.4 Stem

(B,3,640,640)
 → Conv2d(3→4, kernel 3, stride 1, pad 1)
 → BatchNorm2d(4) → SiLU
 = (B,4,640,640)

The stem has 120 parameters. SiLU is x·sigmoid(x). There is no pooling or tokenization.

11.5 Four downsampling and recurrent stages

Each stage performs 3×3 stride-2 convolution, BatchNorm, SiLU, then ConvLSTM. There are no residual blocks, attention layers, transformer tokens, or active sensor-conditioning layers.

Stage Down mapping Feature and state Dilation Down params ConvLSTM params Head input?
1 4→4 B×4×320×320 1 156 1,168 No
2 4→8 B×8×160×160 2 312 4,640 Yes, stride 4
3 8→16 B×16×80×80 4 1,200 18,496 Yes, stride 8
4 16→32 B×32×40×40 8 4,704 73,856 Yes, stride 16

Stage 1 still matters: its high-resolution recurrent output feeds all later stages. The last three tensors are independent multi-scale features. Unlike a classical FPN, there is no top-down pathway, lateral addition, or cross-scale fusion.

11.6 Every ConvLSTM operation

At time t, current feature X_t and previous hidden state H_(t−1) are concatenated. One dilated 3×3 convolution emits four channel groups:

[I,F,O,G] = Conv3×3_dilated([X_t,H_(t−1)])
i=sigmoid(I), f=sigmoid(F), o=sigmoid(O), g=tanh(G)
C_t = f ⊙ C_(t−1) + i ⊙ g
H_t = o ⊙ tanh(C_t)

The input, forget, and output gates govern writing, retaining, and exposing memory. Both H_t and C_t persist. At inference the four pairs have shapes 1×4×320×320, 1×8×160×160, 1×16×80×80, and 1×32×40×40. Dilation gives effective 3×3, 5×5, 9×9, and 17×17 footprints without adding weights. State begins at zero and resets only at sequence boundaries.

11.7 Three active detection heads

For each returned feature, the active regression/objectness tower is:

feature C×H×W
 → Conv2d(C→16,3×3,pad1) → SiLU
 → Conv2d(16→16,3×3,pad1) → SiLU
 ├→ Conv2d(16→4,1×1): raw box vector
 └→ Conv2d(16→1,1×1): objectness logit
Scale Feature Tower + box + objectness params Outputs
stride 4 8×160×160 3,488 + 68 + 17 4×160×160 and 1×160×160
stride 8 16×80×80 4,640 + 68 + 17 4×80×80 and 1×80×80
stride 16 32×40×40 6,944 + 68 + 17 4×40×40 and 1×40×40

There are 25,600 + 6,400 + 1,600 = 33,600 candidate cells. Four regression channels encode raw x/y offsets and log width/height. Objectness alone asks whether an RSO occupies the cell.

11.8 Compatibility classification modules

Three classification towers and 1×1 predictors remain in the checkpoint so historical state-dictionary keys load strictly. They contain 15,123 parameters, but compute_cls=False: their convolutions are not executed, their logits are not trained by the final loss, and the decoder never reads them. In a single-class problem, multiplying objectness by a redundant class probability would square/compress confidence without adding information. Final score is sigmoid(objectness). No quality head exists.

11.9 Auxiliary segmentation head

A nine-parameter 1×1 convolution maps the eight-channel stride-4 feature to one 160×160 logit plane. Per-event labels supervise it with BCE at weight 0.5, helping the finest features retain target-local structure. It runs only while model.training is true; deployment skips it entirely.

11.10 Target assignment and hard loss

A human box is assigned to the scale whose 4×stride nominal size is closest to its maximum dimension. The rounded center and radius-1 neighborhood—up to 3×3 cells—are positive. At cell (g_x,g_y):

t_x=c_x/s−g_x,  t_y=c_y/s−g_y,
t_w=w/s,        t_h=h/s

Objectness uses focal BCE with alpha 0.25 and gamma 2.0, normalized by positive assigned cells and clamped to one on empty windows. The positive objectness floor was 1.0 during final distillation. Regression uses pixel-inclusive CIoU: 1−IoU, normalized center distance, and aspect-ratio penalty. It therefore supplies gradient even before two tiny boxes overlap. Auxiliary BCE and the selective KD terms described in Section 10 complete the objective.

11.11 Anchor-free decode

For raw (r_x,r_y,r_w,r_h) at grid cell (g_x,g_y) and stride s:

c_x=(g_x+r_x)s       c_y=(g_y+r_y)s
w=exp(clamp(r_w,6))s h=exp(clamp(r_h,6))s
score=sigmoid(objectness)

Offsets are intentionally raw, matching training. Exponentiation guarantees positive sizes and the upper clamp prevents overflow.

11.12 Exact stable top-one decode

Production requests one proposal. The decoder concatenates three score maps in fixed scale/grid order, masks values below 0.10, and applies argmax. Exact ties select the first index deterministically. Only the winning cell's box is decoded.

This exactly matches the first result of score-ordered NMS: the global maximum is selected before any overlap comparison can suppress it. Thus production avoids materializing/sorting 33,600 boxes and the Python NMS loop without approximating or changing the selected prediction. Generic pixel-inclusive NMS and its 1,000-candidate pre-bound remain only for multi-detection diagnostics.

11.13 Native-coordinate reconstruction

The winning box is inverted through recorded area-resize geometry with pixel-center rounding. DVX width/height receives the frozen 0.98 calibration before association. Geometry is clipped in native sensor coordinates, but remains floating point until the writer performs final legal integer projection.

11.14 Complete parameter inventory

Module group Parameters Deployment execution
Stem 120 Yes
Four downsampling blocks 6,372 Yes
Four ConvLSTM cells 98,160 Yes
Three regression towers 15,072 Yes
Box predictors 204 Yes
Objectness predictors 51 Yes
Compatibility classification modules 15,123 No
Auxiliary segmentation 9 No
Total 135,111 —

11.15 End-to-end state machine

For each sequence: identify sensor and reset state; slice 40 ms; denoise; form the 3×640×640 tensor; run stem and four ConvLSTMs; run three active heads; choose/decode global top-one; map/calibrate native geometry; update Kalman/coast state; retain at most one box; apply the sensor gate; project integer fields; write the row; and reset everything before another sequence. Docker latency covers this path, not merely the neural forward call.

12. Detection to tracking

The detector proposes appearance-based boxes. The tracker imposes short-horizon physical consistency. Each track maintains:

[centre_x, centre_y, width, height, velocity_x, velocity_y]

Before the next detection arrives, a constant-velocity Kalman model predicts the new center and increases its uncertainty. If a detector box overlaps sufficiently, the state is updated. If evidence disappears briefly, the track may coast for a bounded number of windows with discounted confidence.

The final student policy uses:

  • detector/high/low threshold 0.10;
  • one proposal before tracking and one output after tracking;
  • IoU association threshold 0.40;
  • min_hits=1, max_age=15 internal lifetime;
  • at most one emitted coasted window;
  • coast confidence multiplied by 0.40;
  • no track IDs in the challenge file.

This is intentionally simple. More elaborate motion gates were tested, but measured object motion was slow enough that consecutive boxes overlap. IoU association was therefore more reliable than a complex motion-only discriminator.

13. Distilled student results

Before the final no-retrain adjustments, the distilled checkpoint produced:

Scope Precision Recall F1 mAP TP FP GT
Leakage-free 0.5722 0.6834 0.6229 0.4608 4,472 3,343 6,544
All four 0.5924 0.6899 0.6375 0.5373 4,882 3,359 7,076

The student nearly matched V5 F1 while substantially improving the local mAP diagnostic, despite being 63.6× smaller than one heavy branch. It did not reach the aspirational F1 ≥ 0.70 and mAP > 0.60 targets; those targets also exceed the leakage-free hybrid teacher's measured values.

The full 37,478-window CPU profile measured:

  • weighted mean 18.55 ms/window;
  • worst-sequence p95 21.97 ms/window;
  • every sequence p95 below 30 ms and 40 ms;
  • rare maximum outlier 152.66 ms.

The last number is why both percentiles and maxima are reported. The system comfortably met the steady-state p95 objective locally, but scheduling outliers still existed in the longer profile.

14. Final no-retrain improvements

With the checkpoint and architecture frozen, only low-cost constants were screened:

Setting Previous student Final student
DVX box-size factor 1.00 0.98
DAVIS emit gate 0.30 0.30
DVX emit gate 0.10 0.10
EVK4 emit gate 0.10 0.24
Tracker IoU association 0.30 0.40
Maximum emitted coast 2 windows 1 window
Coast confidence decay 0.50 0.40

Final measured accuracy:

Scope Precision Recall F1 mAP Change in F1 Change in mAP
Leakage-free 0.5796 0.6881 0.6292 0.4645 +0.0063 +0.0037
All four 0.5998 0.6946 0.6437 0.5410 +0.0063 +0.0037

Previous and final student scores

The later 3,920-window profile measured 17.05 ms weighted mean, 20.64 ms worst-sequence p95, and 39.17 ms worst observed window. This shorter profile is not an apples-to-apples latency improvement claim against the earlier 37,478-window run; it confirms that the final policy remained comfortably under the local p95 limit and introduced no expensive operation.

The gains are modest but consistent across both aggregate views. No weight, layer, or tensor shape changed.

15. Current per-sequence behavior

Recording Sensor F1 AP Precision Recall Main interpretation
EVK4 magnitude 7.3 EVK4 0.7647 0.5927 0.7301 0.8027 Strong benefit from EVK4-aware training behavior
DAVIS SAOCOM-1B* DAVIS 0.8610 0.7708 0.9694 0.7744 Strong but session-overlapping
DVX Stars3 DVX 0.6014 0.4414 0.5419 0.6755 Dominant false-positive source
DVX Thuraya3 DVX 0.5166 0.3593 0.7569 0.3921 Dimmest, recall-limited regime

* SAOCOM-1B shares session context with training data and is excluded from the leakage-free aggregate.

DVX remains the main unresolved accuracy problem. Stars3 contributes most false positives, while Thuraya3 has too little signal for adequate recall. These are not equivalent failures: one is primarily ranking/clutter, and the other is primarily missing evidence.

Representative correct detections

Representative localization failures

15.1 Animated detector-and-tracker demonstrations

Four real-data animations accompany this review. Red is the raw top-one detector proposal, cyan is the box emitted after Kalman smoothing, and dashed yellow is human ground truth. The red and cyan trails show recent detector and tracker centers. Each frame also reports confidence, detector IoU, tracker IoU, event count, track ID/age, and elapsed time.

EVK4 detector and tracker animation poster

DAVIS SAOCOM-1B animation poster

DVX Stars3 animation poster

DVX Thuraya3 low-SNR animation poster

These clips were deliberately selected to explain behavior. Their per-frame counts are qualitative descriptions of the shown intervals, not new aggregate accuracy estimates or blind evidence. The animation manifest records the exact sequence/window ranges, promoted cache, final policy, and GIF hashes needed to reproduce them.

16. Docker and reproducibility architecture

The final Linux/amd64 image is designed to be inspected and run without network access. It contains:

  • the exact student checkpoint and SHA-256 sidecar;
  • the compact model definition and preprocessing code;
  • sensor parameters, tracker policy, and writer implementation;
  • pinned CPU runtime dependencies;
  • the official evaluator;
  • an automatic CMD ["sh", "run.sh"] entrypoint;
  • a model-structure document checked against the baked runtime settings.

At runtime it:

  1. reads *.npy sequences from /OrbitSight_dataset;
  2. refuses ambiguous sensors or duplicate sequence names;
  3. verifies the checkpoint hash;
  4. prewarms Numba and one model window;
  5. resets all state at each sequence boundary;
  6. writes <sequencename>.txt files with exactly nine tab-separated columns;
  7. runs the bundled evaluator;
  8. writes Evaluation_Metrics.xlsx;
  9. publishes only the known files under /work/SIMER/DDMMYYYY.

The final improved archive SHA-256 is:

3b2db2bdf3cc7b1d331f8fa6dc8dcc23281bb831177ff07c75f273eb0b819eb5

Archive validation checks the Linux/amd64 manifest, layer diff IDs, automatic command, mount contract, output schema, model hash, runtime dependencies, calibration constants, tracker settings, and baked model-structure text.

17. Evidence boundaries and scientific honesty

Three different claims must not be confused:

  1. Reproducibility: the same files and code reproduce a measured local result.
  2. Score maximization: a setting scores better on recordings already examined.
  3. Blind generalization: a frozen setting performs better on recordings that had no role in training, tuning, diagnostics, or selection.

OrbitSight has strong evidence for the first claim. The routed hybrid and final calibration also have evidence for the second. There are no unused local recordings, so neither provides a new blind confirmation. The organizer's unseen finalist dataset is the correct test of the third claim.

The leakage-free aggregate removes the known DAVIS session overlap, but the remaining files have still been used for diagnostics. “Leakage-free” here means free of that known training- session contamination, not untouched by all development decisions.

18. Why the final student is the recommended submission

The heavy hybrid has the best local accuracy but fails the CPU requirement by an order of magnitude. V5 has a cleaner scientific history but its corrected full-tail latency exceeded 40 ms on some held-out regions, and its local mAP is below the compact student's. The compact student offers:

  • one small checkpoint rather than a routed pair;
  • 135,111 parameters;
  • full 640×640 small-target localization;
  • recurrent temporal integration;
  • sensor-aware but computationally trivial calibration;
  • a local worst-sequence p95 near 21 ms;
  • F1 near the V5 scientific baseline;
  • better local mAP than V5;
  • a smaller, clearer, CPU-only Docker delivery.

This is a practical systems decision rather than a claim that the student is universally more accurate than every predecessor.

19. Recommended next work

If new recordings become available, the next work should be data-centered rather than another threshold sweep:

  1. Freeze the current and rollback Docker hashes before receiving new data.
  2. Keep new labels hidden while both versions produce predictions.
  3. Compare both on exactly the same capture-disjoint recordings.
  4. Require improved F1, no mAP regression, no sensor-family collapse, and p95 below 40 ms.
  5. Treat that dataset as spent after results are opened.

For model research, the highest-value directions are:

  • collect more low-SNR DVX positives and representative empty star fields;
  • train a better proposal-ranking objective for hard negatives;
  • explore calibrated simulation only after matching real event statistics;
  • profile OpenVINO/BF16/INT8 only with strict output-parity and accuracy gates;
  • measure container startup, file reading, formatting, workbook creation, and disk output on the organizer's exact i9-12900H hardware.

20. Conclusion

OrbitSight progressed from an honest sub-0.1 baseline to a complete real-time CPU system by treating data contracts, localization, temporal reasoning, deployment, and evaluation as one problem. V5 established a credible scientific detector. V7 contributed a superior EVK4 specialist. Their hybrid demonstrated the best available local accuracy but exposed an unavoidable latency gap. Selective teacher–student distillation transferred much of that behavior into a model small enough for real-time CPU inference. Final calibration improved the frozen student's operating point without retraining.

The current submission is therefore the product of the whole development path—not merely the last threshold and tracker adjustment. Its strongest qualities are balanced performance, transparent evidence, explicit limitations, and a reproducible offline delivery contract.

Selected references

  1. G. Gallego et al., “Event-based Vision: A Survey,” IEEE TPAMI, 2020. DOI: 10.1109/TPAMI.2020.3008413.
  2. S. Afshar et al., “Event-based Object Detection and Tracking for Space Situational Awareness,” arXiv:1911.08730, 2019.
  3. N. O. Ralph et al., “Astrometric Calibration and Source Characterisation of the Latest Generation Neuromorphic Event-based Cameras for Space Imaging,” arXiv:2211.09939, 2022.
  4. Y. Alkendi et al., “Dynamic Space Object Detection with Neuromorphic Vision Sensors,” European Conference on Space Debris, 2025.
  5. E. Perot et al., “Learning to Detect Objects with a 1 Megapixel Event Camera,” NeurIPS, 2020.
  6. M. Gehrig and D. Scaramuzza, “Recurrent Vision Transformers for Object Detection with Event Cameras,” CVPR, 2023.
  7. X. Shi et al., “Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting,” NeurIPS, 2015.
  8. Z. Tian et al., “FCOS: Fully Convolutional One-Stage Object Detection,” ICCV, 2019.
  9. Z. Ge et al., “YOLOX: Exceeding YOLO Series in 2021,” arXiv:2107.08430, 2021.
  10. T.-Y. Lin et al., “Focal Loss for Dense Object Detection,” ICCV, 2017.
  11. Z. Zheng et al., “Distance-IoU Loss: Faster and Better Learning for Bounding Box Regression,” AAAI, 2020. DOI: 10.1609/aaai.v34i07.6999.
  12. G. Hinton, O. Vinyals, and J. Dean, “Distilling the Knowledge in a Neural Network,” arXiv:1503.02531, 2015.
  13. Y. Zhang et al., “ByteTrack: Multi-Object Tracking by Associating Every Detection Box,” ECCV, 2022.

Primary local evidence

  • docs/technical_report.md
  • docs/v7_audit_status.md
  • notes.md
  • artifacts/v7_gates/official_test_confirmation/epoch33-paired-official-once-v3/report.json
  • artifacts/cpu_student/hybrid_distilled_student_eval_latency_summary.json
  • artifacts/cpu_student/hybrid_distilled_student_full_cpu_latency.json
  • artifacts/cpu_student/post_training_latency_neutral/promotion.json
  • artifacts/cpu_student/post_training_latency_neutral/tracker_live_accuracy.json
  • artifacts/cpu_student/post_training_latency_neutral/tracker_live_latency.json