Download docs/OrbitSight_Full_Technical_Proposal_Review.md from simerai/orbitsight: direct link, hf CLI and curl.
- Browser
- Download file 53.5 kB
-
https://huggingface.co/simerai/orbitsight/resolve/main/docs/OrbitSight_Full_Technical_Proposal_Review.md
- Command line
-
hf download hf://simerai/orbitsight/docs/OrbitSight_Full_Technical_Proposal_Review.md
-
curl -L -o OrbitSight_Full_Technical_Proposal_Review.md https://huggingface.co/simerai/orbitsight/resolve/main/docs/OrbitSight_Full_Technical_Proposal_Review.md
OrbitSight — Full Technical Proposal and Engineering Review
Long-form review draft, 29 August 2026
Team SIMER — Hishaam A., Simreen Siraj, and Christian Sabatini — Ampiere Labs
This is deliberately not the five-page submission proposal. It is the detailed source document from which a shorter competition proposal, pitch deck, or technical appendix can be produced. It tells the complete development story, including failed models, invalid experiments, latency failures, and the limits of the available evaluation evidence.
Executive summary
OrbitSight detects and tracks resident space objects (RSOs) directly from asynchronous
neuromorphic vision sensor events. The final submission model is a compact recurrent neural
network followed by a physically interpretable Kalman tracker. It reads raw *.npy event
recordings, processes consecutive 40 ms windows, and writes the exact ChallengeON prediction
files and evaluation workbook from an automatic, offline, CPU-only Docker container.
The important story is not a single final score. It is a sequence of engineering discoveries:
- The first honest pipeline scored only F1 0.036 / mAP 0.010. Early capped experiments had silently missed most labelled windows, and several apparently good measurements were not reproducible.
- Correct windowing, event denoising, a three-channel representation, recurrent detection, CIoU localization, longer temporal training, and honest best-epoch selection produced V5. After output calibration and gate corrections, V5 reached approximately F1 0.630 / mAP 0.380 on the leakage-free local diagnostic.
- V7 introduced a more principled area-resize and coordinate-mapping contract. It did not beat V5 overall, but it was substantially better on EVK4. V5 remained better on DVX.
- A sensor router combined the specialists: EVK4 → V7, DAVIS/DVX → V5. This hybrid reached F1 0.668 / mAP 0.490 leakage-free and 0.676 / 0.549 across all four local recordings. The route was chosen after per-sensor results were visible, so it is a score-maximizing diagnostic rather than blind model-selection evidence.
- The heavy hybrid was unusable on the competition CPU: approximately 383 ms mean / 451 ms p95, around an order of magnitude above the 40 ms requirement.
- The hybrid therefore became a teacher rather than a deployable model. A 135,111-parameter recurrent student was trained from human labels plus carefully filtered teacher evidence. It retained full 640×640 localization while being approximately 63.6× smaller than one 8,588,339-parameter teacher branch.
- The distilled student measured F1 0.6229 / mAP 0.4608 leakage-free and 0.6375 / 0.5373 on all four recordings, with 18.55 ms mean / 21.97 ms worst-sequence p95 in a full local CPU profile.
- Without retraining, conservative sensor calibration and tracker changes raised the final diagnostics to F1 0.6292 / mAP 0.4645 leakage-free and 0.6437 / 0.5410 all-four. A later 3,920-window profile measured 17.05 ms mean / 20.64 ms worst-sequence p95, with the largest observed window at 39.17 ms.
The result is not the highest-scoring network we built. It is the strongest practical balance of detection quality, CPU latency, reproducibility, and challenge-format correctness.
1. Problem and evaluation objective
1.1 What the sensor produces
A conventional camera exposes a dense image at fixed intervals. A neuromorphic vision sensor works differently: each pixel emits an event when local log-brightness changes enough. An event carries pixel position, polarity, and a microsecond timestamp. The result is a sparse, asynchronous stream rather than a sequence of ordinary photographs.
This is attractive for space situational awareness because event sensors offer high dynamic range, low motion blur, and fine timing. It is also difficult because most events are not the target. In the supplied recordings, labelled RSO events represent roughly 0.08–0.96% of the stream. The rest includes sensor activity, hot pixels, star-field responses, and other background structure.
OrbitSight supports three sensor families with very different native resolutions:
| Sensor | Native resolution | Transformation challenge |
|---|---|---|
| DAVIS346c | 346×260 | Must be enlarged while preserving tiny target geometry |
| DVXplorer | 640×480 | Near the model scale, but strongly affected by low-SNR clutter |
| EVK4 | 1280×720 | Must be reduced without deleting sparse events |
1.2 Why the targets are unusually difficult
- The median labelled box is only about 10×11 pixels; the tenth percentile is about 6×6 pixels.
- IoU ≥ 0.5 on a ten-pixel object allows only a few pixels of center error.
- Signal strength varies by roughly 200× across recordings.
- Every annotated window contains one target box, so a second emitted detection cannot become another true positive; it can only reduce precision.
- The hardest low-SNR cases contain only about six useful object events in a 40 ms window. At that point, localization is limited by sampling noise as much as by model capacity.
1.3 What must be delivered
For every input sequence, the submission produces <sequencename>.txt with these fields:
sequence_id, window_start_timestamp_us, window_end_timestamp_us, x_centre, y_centre, w, h, class_id, confidence
The container also produces Evaluation_Metrics.xlsx. Accuracy is reported using precision,
recall, F1, and AP/mAP at IoU 0.5. The competition additionally calls for end-to-end latency
below 40 ms on an Intel Core i9-12900H CPU. There is no track-ID column, so tracking matters
only when it improves the emitted detection stream.
Literature review and design lineage
OrbitSight is an application-specific synthesis rather than a direct reproduction of one published network. Its design draws on neuromorphic sensing, recurrent event detection, anchor-free dense detection, model compression, and detection-assisted tracking.
Event cameras and space situational awareness
Gallego et al. formalize the event-camera model: a pixel reports a change in log intensity rather than an absolute image sample. The literature emphasizes microsecond timing, high dynamic range, low motion blur, and sparse output, while warning that frame algorithms require an event representation or event-native computation. OrbitSight preserves polarity and within-window timing, but packs events into 40 ms tensors so convolutional and recurrent operations run efficiently on a CPU. Gallego et al., Event-based Vision: A Survey, TPAMI 2020.
Afshar et al. demonstrated event-based RSO detection and tracking across sensors and observing sites. Ralph et al. subsequently studied astrometric calibration and source characterization for event-based space imaging. These works motivate OrbitSight's explicit sensor profiles, native-coordinate reconstruction, and treatment of preprocessing and coordinate mapping as part of the model contract. Afshar et al., Event-based Object Detection and Tracking for Space Situational Awareness, 2019; Ralph et al., Astrometric Calibration and Source Characterisation of Neuromorphic Event-based Cameras for Space Imaging, 2022.
Recurrent event-based object detection
Perot et al.'s Recurrent Event-camera Detector showed that direct event representations combined with recurrent memory outperform feed-forward event detectors because state integrates evidence through time. RVT later showed that multi-stage recurrent backbones can obtain a favorable latency/accuracy tradeoff. OrbitSight adopts direct event tensors, multi-scale features, and persistent temporal state, but uses narrow convolutions and ConvLSTM instead of a transformer for an offline CPU target and tiny objects. Perot et al., Learning to Detect Objects with a 1 Megapixel Event Camera, NeurIPS 2020; Gehrig and Scaramuzza, Recurrent Vision Transformers for Object Detection With Event Cameras, CVPR 2023.
ConvLSTM replaces fully connected state transitions with spatial convolutions, retaining where weak evidence occurred. OrbitSight places one after every downsampling stage and uses dilations 1, 2, 4, and 8 to enlarge spatial context without changing tensor shapes. Shi et al., Convolutional LSTM Network, NeurIPS 2015.
Anchor-free detection, imbalance, and localization
FCOS established fully convolutional anchor-free prediction; YOLOX paired anchor-free boxes with a decoupled head. OrbitSight adopts those ideas with custom assignment and scoring for one RSO class. It predicts raw center offsets and log sizes on three scales without an anchor bank. Tian et al., FCOS, ICCV 2019; Ge et al., YOLOX, 2021.
Dense event tensors create tens of thousands of background cells for at most one target. Focal loss suppresses easy negatives so hard examples control learning. OrbitSight uses sigmoid focal BCE for objectness. CIoU combines overlap, center distance, and aspect-ratio consistency, aligning training more closely with the IoU≥0.5 metric than independent coordinate L1. Lin et al., Focal Loss for Dense Object Detection, ICCV 2017; Zheng et al., Distance-IoU Loss, AAAI 2020.
Distillation and tracking
Knowledge distillation transfers useful behavior from an expensive model or ensemble into a smaller deployable network. OrbitSight follows that deployment logic but does not blindly copy teacher output: human boxes remain authoritative, matched teacher boxes provide low-weight targets, and proposals in human-empty windows become negative evidence. Hinton, Vinyals, and Dean, Distilling the Knowledge in a Neural Network, 2015.
ByteTrack showed the value of associating lower-confidence detections. OrbitSight evaluated that two-tier idea, but production emits only stable global top-one at 0.10, making the nominal low tier unreachable. The final system is therefore a single-tier constant-velocity Kalman/coast tracker using the same association machinery. Zhang et al., ByteTrack, ECCV 2022.
OrbitSight's design contribution
- a polarity/recency representation with sensor-specific event denoising;
- a full-resolution but ultra-narrow recurrent detector;
- single-class objectness decoding without a redundant class multiplier;
- selective, human-authoritative distillation from a sensor-routed teacher;
- weights bound to preprocessing and coordinate-mapping fingerprints;
- exact stable top-one decoding that avoids NMS without changing the selected box;
- one evaluated detector/tracker/writer deployment contract.
2. Development history at a glance
The following table separates comparable leakage-free results from historical development
measurements. The detector_e30 row came from the validation split used at that time and is
therefore useful as progress evidence, not a strict comparison with later test rows.
| Stage | Main change | Leakage-free F1 | Leakage-free mAP | Decision |
|---|---|---|---|---|
| First honest baseline | Correctly measured initial pipeline | 0.036 | 0.010 | Foundation only |
detector_e30 |
30-epoch trained pipeline and tuned tracking | 0.515* | 0.220* | Promising, not directly comparable |
detector_ciou |
CIoU, dilations, regime balancing | 0.5620 | 0.2536 | Retained ideas |
| V5 detector | TBPTT-40, honest split, best epoch | 0.5808 | 0.2749 | Strong detector checkpoint |
| V5 + box calibration | Per-sensor size correction | 0.5933 | 0.2947 | Retained |
| V5 complete pipeline | Low detector gate and per-sensor emit gates | ≈0.6300 | ≈0.3798 | Scientific baseline |
| V6 | Intended area-resize successor | 0.6116† | 0.3570† | Quarantined: invalid contract |
| V7 epoch 33 | Correct area resize and pixel-center mapping | 0.5954 | 0.3585 | Lost overall; EVK4 specialist |
| Routed V5/V7 hybrid | EVK4 V7, DAVIS/DVX V5 | 0.6683 | 0.4895 | Accuracy teacher; too slow |
| Compact student from scratch | Same small architecture, no teacher | 0.1055 aggregate | 0.2602 aggregate | DVX collapsed |
| Distilled CPU student | Human labels + selective hybrid teaching | 0.6229 | 0.4608 | Fast deployable candidate |
| Final CPU student | No-retrain calibration and tracker refinement | 0.6292 | 0.4645 | Recommended submission |
* Historical validation regime.
† Retained only as an audit trail because training inputs and declared preprocessing did not match.
Two lessons are visible immediately. First, the large gain from below 0.1 to approximately 0.63 did not come from one architectural trick; it came from correcting the entire data, training, inference, tracking, and writing contract. Second, the best accuracy model was not the best deployment model. The heavy hybrid forced a separate compression phase.
3. Why the first scores were below 0.1
3.1 The first honest result
The first reproducible leakage-free baseline was F1 0.0364 / mAP 0.0096. This was useful because it established an honest floor, even though the score was extremely low. An earlier F1 value around 0.14 came from a measurement path that could not be reproduced consistently and is not used as the project baseline.
3.2 The label-starvation bug
Early capped experiments evaluated the first N windows of a stream. Labels, however, often began much later. In one sequence the first labelled window was 1,562, so a 1,500-window cap could finish before reaching the main annotated region. Across the affected capped runs, only about 34% of labels were visible.
This bug mattered twice: training batches contained too few positive examples, and evaluation looked at the wrong temporal region. The fix introduced an explicit starting window and aligned capped runs to the first labelled window. This was not a model improvement; it was a measurement correction that made later model improvements meaningful.
3.3 Other foundational corrections
- Ground-truth timestamps were required to lie on the exact 40 ms window phase.
- The official evaluator's pixel-inclusive IoU convention was reproduced locally.
- Prediction box fields were made integer-valued only at the final writer boundary.
- Zero-sized boxes after rounding were clamped to at least one pixel.
- Model objectness was decoded correctly for the single-class RSO task instead of being multiplied by a redundant class probability.
- Detector, recurrent, and tracker state were reset at every sequence boundary.
These corrections are easy to underestimate. A sophisticated model evaluated through a misaligned or inconsistent pipeline is still a bad system.
4. The common event-processing pipeline
The heavy models and compact student share the same conceptual flow:
raw event stream
↓
memory-mapped ingestion and phase-checked 40 ms windows
↓
sensor-specific background-activity filtering
↓
three-channel event tensor at 640×640
↓
recurrent anchor-free detector
↓
map box back to native sensor coordinates
↓
Kalman association and bounded coasting
↓
sensor-specific output gate and one-box cap
↓
integer portal row + Evaluation_Metrics.xlsx
4.1 Background-activity filtering
A real optical signal tends to produce neighboring events close together in time. Independent sensor noise is less likely to have that support. The BA filter asks a simple question for each event: did enough nearby pixels fire recently? The parameters are sensor-specific:
| Sensor | Required support | Radius | Time horizon | RSO events retained | Noise rejected |
|---|---|---|---|---|---|
| DAVIS | 1 | 6 px | 40 ms | 99.47% | 28.32% |
| DVX | 1 | 4 px | 20 ms | 98.24% | 39.94% |
| EVK4 | 1 | 3 px | 5 ms | 97.71% | 41.41% |
The filter is window-local and stateless in the deployed route. Hot-pixel support exists but is not silently enabled because changing preprocessing only at inference would create a train/deployment mismatch.
4.2 Three-channel representation
Each 40 ms event window becomes a three-channel tensor:
- positive-event count;
- negative-event count;
- recency of the latest event.
The first two channels show where brightness increased or decreased. The recency channel preserves temporal direction inside the window. A moving point source often produces a polarity pattern and timestamp gradient, so these three channels retain more useful physics than a single accumulated image.
The model input remains 640×640 even in the compact student. Reducing the image would save compute, but a 6–11 pixel target cannot tolerate much spatial quantization. We reduced network width instead of discarding localization resolution.
5. From early trained models to V5
5.1 The detector family
OrbitSight uses an anchor-free, multi-scale detector. “Anchor-free” means that the network predicts a target center and size directly at grid locations instead of choosing from a bank of predefined rectangle shapes. This suits small RSOs whose apparent extent depends on brightness, motion, optics, and sensor resolution.
The backbone contains four downsampling stages. Each stage includes ConvLSTM recurrence, so
the model carries information from earlier windows. Recurrent dilations (1, 2, 4, 8) expand
the spatial context without requiring a much deeper network. Detection heads operate at
multiple scales, with stride 4 retained for precise small-target localization.
5.2 CIoU and confidence learning
Early box regression did not optimize the competition's overlap criterion closely enough. Complete IoU (CIoU) improved the learning signal by combining overlap, center distance, and shape consistency. It also supplies a gradient when a predicted box and target do not yet overlap.
The objectness target is linked to achieved localization quality. In plain language: a box should be confident when it is both likely to contain the RSO and accurately localized. This makes confidence useful for ranking, but it also means correct tiny boxes may have moderate rather than near-one scores. That observation later motivated separating the detector's low proposal threshold from the final output gate.
5.3 Training improvements that produced V5
- Regime balancing: batches were designed to prevent abundant background windows from erasing rare positive examples.
- TBPTT length 40: the network learned across 1.6 seconds of 40 ms windows instead of a shorter temporal horizon.
- Honest validation: sequence-level separation reduced leakage from adjacent windows.
- Best-epoch selection: V5 kept epoch 33 rather than simply shipping the final epoch.
- Full-path validation: evaluation included tracking, coasting, gates, and writer quantization—the path that would actually be submitted.
V5's heavy backbone widths are (32, 64, 128, 256), the head width is 128, and the model has
8,588,339 parameters. It became the scientific baseline because it was the strongest
model supported by a coherent, reproducible contract.
5.4 The output-gate breakthrough
Originally, the detector threshold and final output threshold were effectively the same. A box below the emit threshold was discarded before tracking, even if it was the best proposal in the window. The corrected design runs the detector at a low 0.10 proposal threshold, lets the tracker use that evidence, and applies the sensor output gate only at the writer.
This separation was one of the largest no-retrain gains in the heavy pipeline. The retained V5 gates were DAVIS 0.30, DVX 0.10, and EVK4 0.10. Per-sensor box-size correction also repaired systematic under-sizing caused by sensor transformations. Together, these changes moved the complete V5 pipeline to approximately F1 0.630 / mAP 0.380 leakage-free.
6. V6 and the importance of a preprocessing contract
V6 was intended to test area resampling and add dim-regime data. Its recorded score appeared competitive, but an audit found a critical mismatch: training consumed historical nearest- resized tensors while metadata claimed area resize. Evaluation could therefore feed one representation to weights trained on another.
The checkpoint still loaded because tensor shapes matched. That is exactly why the problem was dangerous: shape compatibility did not imply semantic compatibility. V6 was quarantined, and later loaders were strengthened to bind checkpoints to preprocessing fingerprints.
The general lesson is central to OrbitSight: a model is not just weights. It is weights plus window definition, denoising, resize method, coordinate mapping, count transform, state behavior, thresholds, and writer projection.
7. V7: better geometry, mixed sensor behavior
V7 rebuilt the preprocessing contract around true area resampling and pixel-center rounding. The intent was sound: when EVK4 is reduced from 1280×720, nearest-neighbor sampling can delete sparse events, while area resampling preserves aggregate activity more faithfully.
The consolidated V7 model did not beat V5 overall:
| Model | Leakage-free F1 | Leakage-free mAP | All-four F1 | All-four mAP |
|---|---|---|---|---|
| Paired V5 | 0.6298 | 0.3804 | 0.6403 | 0.4670 |
| V7 epoch 33 | 0.5954 | 0.3585 | 0.6100 | 0.4378 |
However, aggregate results hid meaningful specialization:
| Sequence | V5 F1 / AP | V7 F1 / AP | Better branch |
|---|---|---|---|
| EVK4 magnitude 7.3 | 0.5175 / 0.2797 | 0.7121 / 0.6070 | V7 by a large margin |
| DVX Stars3 | 0.6667 / 0.5812 | 0.5723 / 0.3525 | V5 |
| DVX Thuraya3 | 0.4357 / 0.2803 | 0.2095 / 0.1160 | V5 |
V7 was therefore not a universal successor. It was an EVK4 specialist. This distinction created the hybrid teacher.
8. The V5/V7 sensor-routed hybrid
The router chooses one heavy branch once per sequence:
- EVK4 → V7 epoch 33;
- DAVIS → V5 best;
- DVX → V5 best.
Only one branch runs for a given window, so this is routing rather than simultaneous ensembling. The routing result was:
| Scope | Precision | Recall | F1 | mAP |
|---|---|---|---|---|
| Leakage-free | 0.6400 | 0.6991 | 0.6683 | 0.4895 |
| All four | 0.6504 | 0.7036 | 0.6760 | 0.5489 |
These were the highest practical local accuracy values in the project. They are also post-hoc: the route was proposed after the per-sensor results were known. The hybrid is therefore useful as a score-maximizing package and as a teacher, but its local result is not an unbiased estimate of unseen performance.
8.1 Why the hybrid could not be submitted on CPU
Each branch still has 8.59 million parameters and full-width recurrent computation. On a local four-thread CPU, the hybrid measured approximately:
- 383 ms mean per window;
- 451 ms p95 per window;
- roughly 313 ms/window effective during the complete four-sequence replay.
The budget is 40 ms. This is not a small optimization gap. Tracker tuning, output formatting, or minor Python changes cannot create a tenfold speedup in the dominant convolutional model. The architecture had to become much smaller.
9. Why teacher–student distillation was chosen
The hybrid knew useful sensor-specific behavior, but was too slow to deploy. Distillation separates those roles:
- the teacher is allowed to be large and expensive because it runs only while generating training evidence;
- the student learns from human labels plus selected teacher information, then runs alone in the submission container.
The shipped image does not contain V5, V7, or the router. Teacher latency is therefore absent from deployment latency.
9.1 Why a small model trained from scratch was not enough
The same compact architecture was first trained without the routed teacher. Its aggregate F1/mAP was 0.1055 / 0.2602 and DVX effectively collapsed. On a six-positive DVX held-out capture, the best examined epoch produced 2 TP, 3,941 FP, and 4 FN. The two true positives ranked 793rd and 2,595th by confidence.
No output threshold can repair that ordering. Raising the threshold removes the buried true positives along with false positives; lowering it emits thousands of false detections. The failure was a hard-negative ranking problem, which is exactly where a stronger teacher can provide useful structure.
9.2 Selective distillation rather than blind imitation
Teacher outputs are not automatically true. The training design keeps human ground truth authoritative:
- A teacher detection that matches human GT may contribute a low-weight soft distillation target.
- A human GT missed by the teacher remains a full positive.
- A teacher detection in a human-empty window is treated as hard-negative/ranking evidence, not copied as a positive.
- Post-tracker teacher boxes are diagnostics only; the student learns pre-tracker detection behavior so it is not tracked twice.
- Teacher recurrence, student recurrence, BA state, and tracking state reset at sequence boundaries.
- Cache manifests bind source hashes, sensor parameters, preprocessing, coordinate mapping, checkpoint identity, and code/schema versions.
This design prevents the most dangerous distillation failure: teaching the student to repeat the teacher's false positives.
9.3 Training evidence and run
Teacher evidence was generated in two independent shards and merged under signed manifests. Across the merged cache there were:
- 17 training sequences;
- 106,194 windows;
- 18,139 pre-tracker teacher detections;
- 16,682 post-tracker diagnostic detections;
- 11,343 positive windows with a matching teacher detection;
- 5,238 human-empty windows containing teacher proposals, retained as hard negatives.
The final A100 training used 36 epochs, 260 updates per epoch, 9,360 optimizer updates, batch 16, and segment length 20. An interrupted client connection was recovered by exact checkpoint resume, including model, optimizer, scheduler, RNG, sampler, and batching state. The final checkpoint SHA-256 is:
5f176b9baffc3478b1aaada13277e2c6093ba02439b76b8330cb8d654f5f554d
10. Final-model methodology
This section describes how the final student was produced and evaluated. Section 11 then defines the frozen network and inference path component by component.
10.1 Research questions and temporal unit
Development asked whether one model could localize tiny RSOs across all three sensor families, inherit useful sensor-specific behavior from the heavy hybrid, and keep the complete deployed path below 40 ms. The atomic example is a half-open 40,000 µs event interval. Window phase is anchored to the sequence timestamp contract. Recurrent, BA, and tracker state are initialized at the start of a sequence, advanced exactly once per window, and destroyed at the boundary.
Training uses ordered segments, not shuffled independent frames. Final distillation used 20 consecutive windows per segment—0.8 seconds of temporal context—and batches of 16 independent segments. Each batch member carries its own four hidden/cell state pairs before truncated backpropagation detaches them.
10.2 Data preparation
For every window the method:
- memory-maps and timestamp-slices events;
- applies the sensor-specific neighborhood/time BA rule;
- builds positive-count, negative-count, and latest-event-recency planes;
- area-resizes the planes to 640×640;
- maps human boxes with the same pixel-center convention;
- builds the optional per-event auxiliary target.
Counts remain raw. Stateful BA, sensor FiLM, hot-pixel changes, and target-event thinning are disabled in the final route. Cache fingerprints bind the source, sensor profile, 40 ms phase, input size, count transform, area resize, and coordinate mapping. This prevents a shape-compatible checkpoint from being used with semantically different preprocessing, the failure that invalidated V6.
10.3 Human-authoritative teacher evidence
The V5/V7 router runs offline. Its pre-tracker proposals are stored in sidecars aligned one-to-one with student windows and marked as matched or unmatched against human GT. Post-tracker teacher boxes are diagnostic only. The selective objective is:
L_total = L_human
+ 0.15 L_teacher-objectness-on-matched-GT
+ 0.25 L_teacher-box-on-matched-GT
+ 0.35 L_teacher-empty-window-negative
Teacher scores below 0.05 are ignored. Matched proposals supply a BCE confidence target and Smooth-L1 encoded-box target. A proposal in a human-empty window pushes the corresponding student objectness toward zero. An unmatched teacher proposal never creates a positive label. The teacher is therefore a ranking guide and hard-negative miner, not a replacement annotator.
10.4 Optimization protocol
The final A100 run used 36 epochs, 260 updates per epoch, and 9,360 optimizer steps:
| Item | Final setting |
|---|---|
| Optimizer | Adam |
| Initial / minimum LR | 5×10⁻⁴ / 1×10⁻⁵ |
| Warm-up / schedule | 100 steps / cosine decay |
| Precision | FP32; TF32 disabled |
| Gradient-norm clip | 1.0 |
| Batch / segment | 16 / 20 windows |
| Negative/positive target | 2.0 |
| Maximum oversampling | 4.0× |
| Human box loss | pixel-inclusive CIoU, weight 2.0 |
| Auxiliary segmentation | BCE, weight 0.5 |
| Objectness target floor | 1.0 throughout |
| Quality head / loss | disabled / 0.0 |
| Seed | 1337 |
Adam state, scheduler, RNG, batching order, partial epoch progress, preprocessing fingerprint, and the training contract were included in exact-resume checkpoints. A remote disconnect therefore resumed the same stochastic run rather than launching a similar but different experiment.
10.5 Evaluation methodology
Predictions are scored only after native mapping, calibration, tracking, clipping, top-one selection, sensor output gating, and integer writer projection. IoU uses the official pixel-inclusive convention. Metrics are reported for the three-sequence leakage-free aggregate and all four local recordings. Because every remaining recording has influenced diagnostics, these are reproducibility and local comparison results—not blind-generalization proof. Latency is measured after warm-up at streaming batch one and reports mean, per-sequence p95, and maximum rather than forward-pass mean alone.
11. Final model architecture: complete component reference
11.1 Frozen architecture contract
| Property | Final value |
|---|---|
| Class | OrbitSightDetector |
| Inference input | 1×3×640×640 |
| Backbone widths | 4, 8, 16, 32 |
| ConvLSTM dilations | 1, 2, 4, 8 |
| Detection scales | stride 4, 8, 16 |
| Head width | 16 |
| Classes | one RSO foreground class |
| Stored parameters | 135,111 |
| Quality head / sensor FiLM | absent / absent |
| Auxiliary segmentation | stored; training-only execution |
| Classification towers | stored for compatibility; never executed |
11.2 Compression decision
| Component | Heavy V5/V7 branch | CPU student |
|---|---|---|
| Input | 3×640×640 | 3×640×640 |
| Backbone widths | 32, 64, 128, 256 | 4, 8, 16, 32 |
| Head width | 128 | 16 |
| Dilations | 1, 2, 4, 8 | 1, 2, 4, 8 |
| Parameters | 8,588,339 | 135,111 |
| Runtime role | Offline teacher | Deployed model |
The 63.6× reduction comes mainly from channel width. Input resolution, recurrence, the finest stride-4 grid, and three detection scales remain because spatial quantization is expensive for 6–11 pixel targets.
11.3 Input representation
X ∈ R^(B×3×640×640) is float32. Channel 0 counts positive-polarity events; channel 1 counts negative
events; channel 2 records normalized recency of the latest event. Empty pixels are zero. This preserves
brightness-change sign and within-window order without reconstructing an intensity image.
11.4 Stem
(B,3,640,640)
→ Conv2d(3→4, kernel 3, stride 1, pad 1)
→ BatchNorm2d(4) → SiLU
= (B,4,640,640)
The stem has 120 parameters. SiLU is x·sigmoid(x). There is no pooling or tokenization.
11.5 Four downsampling and recurrent stages
Each stage performs 3×3 stride-2 convolution, BatchNorm, SiLU, then ConvLSTM. There are no residual blocks, attention layers, transformer tokens, or active sensor-conditioning layers.
| Stage | Down mapping | Feature and state | Dilation | Down params | ConvLSTM params | Head input? |
|---|---|---|---|---|---|---|
| 1 | 4→4 | B×4×320×320 | 1 | 156 | 1,168 | No |
| 2 | 4→8 | B×8×160×160 | 2 | 312 | 4,640 | Yes, stride 4 |
| 3 | 8→16 | B×16×80×80 | 4 | 1,200 | 18,496 | Yes, stride 8 |
| 4 | 16→32 | B×32×40×40 | 8 | 4,704 | 73,856 | Yes, stride 16 |
Stage 1 still matters: its high-resolution recurrent output feeds all later stages. The last three tensors are independent multi-scale features. Unlike a classical FPN, there is no top-down pathway, lateral addition, or cross-scale fusion.
11.6 Every ConvLSTM operation
At time t, current feature X_t and previous hidden state H_(t−1) are concatenated. One dilated
3×3 convolution emits four channel groups:
[I,F,O,G] = Conv3×3_dilated([X_t,H_(t−1)])
i=sigmoid(I), f=sigmoid(F), o=sigmoid(O), g=tanh(G)
C_t = f ⊙ C_(t−1) + i ⊙ g
H_t = o ⊙ tanh(C_t)
The input, forget, and output gates govern writing, retaining, and exposing memory. Both H_t and
C_t persist. At inference the four pairs have shapes 1×4×320×320, 1×8×160×160,
1×16×80×80, and 1×32×40×40. Dilation gives effective 3×3, 5×5, 9×9, and 17×17 footprints without
adding weights. State begins at zero and resets only at sequence boundaries.
11.7 Three active detection heads
For each returned feature, the active regression/objectness tower is:
feature C×H×W
→ Conv2d(C→16,3×3,pad1) → SiLU
→ Conv2d(16→16,3×3,pad1) → SiLU
├→ Conv2d(16→4,1×1): raw box vector
└→ Conv2d(16→1,1×1): objectness logit
| Scale | Feature | Tower + box + objectness params | Outputs |
|---|---|---|---|
| stride 4 | 8×160×160 | 3,488 + 68 + 17 | 4×160×160 and 1×160×160 |
| stride 8 | 16×80×80 | 4,640 + 68 + 17 | 4×80×80 and 1×80×80 |
| stride 16 | 32×40×40 | 6,944 + 68 + 17 | 4×40×40 and 1×40×40 |
There are 25,600 + 6,400 + 1,600 = 33,600 candidate cells. Four regression channels encode raw x/y offsets and log width/height. Objectness alone asks whether an RSO occupies the cell.
11.8 Compatibility classification modules
Three classification towers and 1×1 predictors remain in the checkpoint so historical state-dictionary
keys load strictly. They contain 15,123 parameters, but compute_cls=False: their convolutions are not
executed, their logits are not trained by the final loss, and the decoder never reads them. In a
single-class problem, multiplying objectness by a redundant class probability would square/compress
confidence without adding information. Final score is sigmoid(objectness). No quality head exists.
11.9 Auxiliary segmentation head
A nine-parameter 1×1 convolution maps the eight-channel stride-4 feature to one 160×160 logit plane.
Per-event labels supervise it with BCE at weight 0.5, helping the finest features retain target-local
structure. It runs only while model.training is true; deployment skips it entirely.
11.10 Target assignment and hard loss
A human box is assigned to the scale whose 4×stride nominal size is closest to its maximum dimension.
The rounded center and radius-1 neighborhood—up to 3×3 cells—are positive. At cell (g_x,g_y):
t_x=c_x/s−g_x, t_y=c_y/s−g_y,
t_w=w/s, t_h=h/s
Objectness uses focal BCE with alpha 0.25 and gamma 2.0, normalized by positive assigned cells and
clamped to one on empty windows. The positive objectness floor was 1.0 during final distillation.
Regression uses pixel-inclusive CIoU: 1−IoU, normalized center distance, and aspect-ratio penalty.
It therefore supplies gradient even before two tiny boxes overlap. Auxiliary BCE and the selective KD
terms described in Section 10 complete the objective.
11.11 Anchor-free decode
For raw (r_x,r_y,r_w,r_h) at grid cell (g_x,g_y) and stride s:
c_x=(g_x+r_x)s c_y=(g_y+r_y)s
w=exp(clamp(r_w,6))s h=exp(clamp(r_h,6))s
score=sigmoid(objectness)
Offsets are intentionally raw, matching training. Exponentiation guarantees positive sizes and the upper clamp prevents overflow.
11.12 Exact stable top-one decode
Production requests one proposal. The decoder concatenates three score maps in fixed scale/grid order,
masks values below 0.10, and applies argmax. Exact ties select the first index deterministically. Only
the winning cell's box is decoded.
This exactly matches the first result of score-ordered NMS: the global maximum is selected before any overlap comparison can suppress it. Thus production avoids materializing/sorting 33,600 boxes and the Python NMS loop without approximating or changing the selected prediction. Generic pixel-inclusive NMS and its 1,000-candidate pre-bound remain only for multi-detection diagnostics.
11.13 Native-coordinate reconstruction
The winning box is inverted through recorded area-resize geometry with pixel-center rounding. DVX width/height receives the frozen 0.98 calibration before association. Geometry is clipped in native sensor coordinates, but remains floating point until the writer performs final legal integer projection.
11.14 Complete parameter inventory
| Module group | Parameters | Deployment execution |
|---|---|---|
| Stem | 120 | Yes |
| Four downsampling blocks | 6,372 | Yes |
| Four ConvLSTM cells | 98,160 | Yes |
| Three regression towers | 15,072 | Yes |
| Box predictors | 204 | Yes |
| Objectness predictors | 51 | Yes |
| Compatibility classification modules | 15,123 | No |
| Auxiliary segmentation | 9 | No |
| Total | 135,111 | — |
11.15 End-to-end state machine
For each sequence: identify sensor and reset state; slice 40 ms; denoise; form the 3×640×640 tensor; run stem and four ConvLSTMs; run three active heads; choose/decode global top-one; map/calibrate native geometry; update Kalman/coast state; retain at most one box; apply the sensor gate; project integer fields; write the row; and reset everything before another sequence. Docker latency covers this path, not merely the neural forward call.
12. Detection to tracking
The detector proposes appearance-based boxes. The tracker imposes short-horizon physical consistency. Each track maintains:
[centre_x, centre_y, width, height, velocity_x, velocity_y]
Before the next detection arrives, a constant-velocity Kalman model predicts the new center and increases its uncertainty. If a detector box overlaps sufficiently, the state is updated. If evidence disappears briefly, the track may coast for a bounded number of windows with discounted confidence.
The final student policy uses:
- detector/high/low threshold 0.10;
- one proposal before tracking and one output after tracking;
- IoU association threshold 0.40;
min_hits=1,max_age=15internal lifetime;- at most one emitted coasted window;
- coast confidence multiplied by 0.40;
- no track IDs in the challenge file.
This is intentionally simple. More elaborate motion gates were tested, but measured object motion was slow enough that consecutive boxes overlap. IoU association was therefore more reliable than a complex motion-only discriminator.
13. Distilled student results
Before the final no-retrain adjustments, the distilled checkpoint produced:
| Scope | Precision | Recall | F1 | mAP | TP | FP | GT |
|---|---|---|---|---|---|---|---|
| Leakage-free | 0.5722 | 0.6834 | 0.6229 | 0.4608 | 4,472 | 3,343 | 6,544 |
| All four | 0.5924 | 0.6899 | 0.6375 | 0.5373 | 4,882 | 3,359 | 7,076 |
The student nearly matched V5 F1 while substantially improving the local mAP diagnostic, despite being 63.6× smaller than one heavy branch. It did not reach the aspirational F1 ≥ 0.70 and mAP > 0.60 targets; those targets also exceed the leakage-free hybrid teacher's measured values.
The full 37,478-window CPU profile measured:
- weighted mean 18.55 ms/window;
- worst-sequence p95 21.97 ms/window;
- every sequence p95 below 30 ms and 40 ms;
- rare maximum outlier 152.66 ms.
The last number is why both percentiles and maxima are reported. The system comfortably met the steady-state p95 objective locally, but scheduling outliers still existed in the longer profile.
14. Final no-retrain improvements
With the checkpoint and architecture frozen, only low-cost constants were screened:
| Setting | Previous student | Final student |
|---|---|---|
| DVX box-size factor | 1.00 | 0.98 |
| DAVIS emit gate | 0.30 | 0.30 |
| DVX emit gate | 0.10 | 0.10 |
| EVK4 emit gate | 0.10 | 0.24 |
| Tracker IoU association | 0.30 | 0.40 |
| Maximum emitted coast | 2 windows | 1 window |
| Coast confidence decay | 0.50 | 0.40 |
Final measured accuracy:
| Scope | Precision | Recall | F1 | mAP | Change in F1 | Change in mAP |
|---|---|---|---|---|---|---|
| Leakage-free | 0.5796 | 0.6881 | 0.6292 | 0.4645 | +0.0063 | +0.0037 |
| All four | 0.5998 | 0.6946 | 0.6437 | 0.5410 | +0.0063 | +0.0037 |
The later 3,920-window profile measured 17.05 ms weighted mean, 20.64 ms worst-sequence p95, and 39.17 ms worst observed window. This shorter profile is not an apples-to-apples latency improvement claim against the earlier 37,478-window run; it confirms that the final policy remained comfortably under the local p95 limit and introduced no expensive operation.
The gains are modest but consistent across both aggregate views. No weight, layer, or tensor shape changed.
15. Current per-sequence behavior
| Recording | Sensor | F1 | AP | Precision | Recall | Main interpretation |
|---|---|---|---|---|---|---|
| EVK4 magnitude 7.3 | EVK4 | 0.7647 | 0.5927 | 0.7301 | 0.8027 | Strong benefit from EVK4-aware training behavior |
| DAVIS SAOCOM-1B* | DAVIS | 0.8610 | 0.7708 | 0.9694 | 0.7744 | Strong but session-overlapping |
| DVX Stars3 | DVX | 0.6014 | 0.4414 | 0.5419 | 0.6755 | Dominant false-positive source |
| DVX Thuraya3 | DVX | 0.5166 | 0.3593 | 0.7569 | 0.3921 | Dimmest, recall-limited regime |
* SAOCOM-1B shares session context with training data and is excluded from the leakage-free aggregate.
DVX remains the main unresolved accuracy problem. Stars3 contributes most false positives, while Thuraya3 has too little signal for adequate recall. These are not equivalent failures: one is primarily ranking/clutter, and the other is primarily missing evidence.
15.1 Animated detector-and-tracker demonstrations
Four real-data animations accompany this review. Red is the raw top-one detector proposal, cyan is the box emitted after Kalman smoothing, and dashed yellow is human ground truth. The red and cyan trails show recent detector and tracker centers. Each frame also reports confidence, detector IoU, tracker IoU, event count, track ID/age, and elapsed time.
- EVK4 bright-target animation: 32 consecutive windows from the magnitude-7.3 EVK4 recording. Tracker IoU exceeded raw detector IoU in 27 of the 32 displayed labelled windows.
- DAVIS SAOCOM-1B tracker-smoothing animation: 29 consecutive windows at native 346×260 resolution. Tracker IoU was higher in 15 of the displayed labelled windows.
- DVX Stars3 tracker-smoothing animation: 30 consecutive cluttered windows. Tracker IoU was higher in 18 displayed windows.
- DVX Thuraya3 low-SNR animation: 33 consecutive windows in the hardest dim-target regime; 30 contain GT, and tracker IoU was higher in 19 of those labelled frames.
These clips were deliberately selected to explain behavior. Their per-frame counts are qualitative descriptions of the shown intervals, not new aggregate accuracy estimates or blind evidence. The animation manifest records the exact sequence/window ranges, promoted cache, final policy, and GIF hashes needed to reproduce them.
16. Docker and reproducibility architecture
The final Linux/amd64 image is designed to be inspected and run without network access. It contains:
- the exact student checkpoint and SHA-256 sidecar;
- the compact model definition and preprocessing code;
- sensor parameters, tracker policy, and writer implementation;
- pinned CPU runtime dependencies;
- the official evaluator;
- an automatic
CMD ["sh", "run.sh"]entrypoint; - a model-structure document checked against the baked runtime settings.
At runtime it:
- reads
*.npysequences from/OrbitSight_dataset; - refuses ambiguous sensors or duplicate sequence names;
- verifies the checkpoint hash;
- prewarms Numba and one model window;
- resets all state at each sequence boundary;
- writes
<sequencename>.txtfiles with exactly nine tab-separated columns; - runs the bundled evaluator;
- writes
Evaluation_Metrics.xlsx; - publishes only the known files under
/work/SIMER/DDMMYYYY.
The final improved archive SHA-256 is:
3b2db2bdf3cc7b1d331f8fa6dc8dcc23281bb831177ff07c75f273eb0b819eb5
Archive validation checks the Linux/amd64 manifest, layer diff IDs, automatic command, mount contract, output schema, model hash, runtime dependencies, calibration constants, tracker settings, and baked model-structure text.
17. Evidence boundaries and scientific honesty
Three different claims must not be confused:
- Reproducibility: the same files and code reproduce a measured local result.
- Score maximization: a setting scores better on recordings already examined.
- Blind generalization: a frozen setting performs better on recordings that had no role in training, tuning, diagnostics, or selection.
OrbitSight has strong evidence for the first claim. The routed hybrid and final calibration also have evidence for the second. There are no unused local recordings, so neither provides a new blind confirmation. The organizer's unseen finalist dataset is the correct test of the third claim.
The leakage-free aggregate removes the known DAVIS session overlap, but the remaining files have still been used for diagnostics. “Leakage-free” here means free of that known training- session contamination, not untouched by all development decisions.
18. Why the final student is the recommended submission
The heavy hybrid has the best local accuracy but fails the CPU requirement by an order of magnitude. V5 has a cleaner scientific history but its corrected full-tail latency exceeded 40 ms on some held-out regions, and its local mAP is below the compact student's. The compact student offers:
- one small checkpoint rather than a routed pair;
- 135,111 parameters;
- full 640×640 small-target localization;
- recurrent temporal integration;
- sensor-aware but computationally trivial calibration;
- a local worst-sequence p95 near 21 ms;
- F1 near the V5 scientific baseline;
- better local mAP than V5;
- a smaller, clearer, CPU-only Docker delivery.
This is a practical systems decision rather than a claim that the student is universally more accurate than every predecessor.
19. Recommended next work
If new recordings become available, the next work should be data-centered rather than another threshold sweep:
- Freeze the current and rollback Docker hashes before receiving new data.
- Keep new labels hidden while both versions produce predictions.
- Compare both on exactly the same capture-disjoint recordings.
- Require improved F1, no mAP regression, no sensor-family collapse, and p95 below 40 ms.
- Treat that dataset as spent after results are opened.
For model research, the highest-value directions are:
- collect more low-SNR DVX positives and representative empty star fields;
- train a better proposal-ranking objective for hard negatives;
- explore calibrated simulation only after matching real event statistics;
- profile OpenVINO/BF16/INT8 only with strict output-parity and accuracy gates;
- measure container startup, file reading, formatting, workbook creation, and disk output on the organizer's exact i9-12900H hardware.
20. Conclusion
OrbitSight progressed from an honest sub-0.1 baseline to a complete real-time CPU system by treating data contracts, localization, temporal reasoning, deployment, and evaluation as one problem. V5 established a credible scientific detector. V7 contributed a superior EVK4 specialist. Their hybrid demonstrated the best available local accuracy but exposed an unavoidable latency gap. Selective teacher–student distillation transferred much of that behavior into a model small enough for real-time CPU inference. Final calibration improved the frozen student's operating point without retraining.
The current submission is therefore the product of the whole development path—not merely the last threshold and tracker adjustment. Its strongest qualities are balanced performance, transparent evidence, explicit limitations, and a reproducible offline delivery contract.
Selected references
- G. Gallego et al., “Event-based Vision: A Survey,” IEEE TPAMI, 2020. DOI: 10.1109/TPAMI.2020.3008413.
- S. Afshar et al., “Event-based Object Detection and Tracking for Space Situational Awareness,” arXiv:1911.08730, 2019.
- N. O. Ralph et al., “Astrometric Calibration and Source Characterisation of the Latest Generation Neuromorphic Event-based Cameras for Space Imaging,” arXiv:2211.09939, 2022.
- Y. Alkendi et al., “Dynamic Space Object Detection with Neuromorphic Vision Sensors,” European Conference on Space Debris, 2025.
- E. Perot et al., “Learning to Detect Objects with a 1 Megapixel Event Camera,” NeurIPS, 2020.
- M. Gehrig and D. Scaramuzza, “Recurrent Vision Transformers for Object Detection with Event Cameras,” CVPR, 2023.
- X. Shi et al., “Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting,” NeurIPS, 2015.
- Z. Tian et al., “FCOS: Fully Convolutional One-Stage Object Detection,” ICCV, 2019.
- Z. Ge et al., “YOLOX: Exceeding YOLO Series in 2021,” arXiv:2107.08430, 2021.
- T.-Y. Lin et al., “Focal Loss for Dense Object Detection,” ICCV, 2017.
- Z. Zheng et al., “Distance-IoU Loss: Faster and Better Learning for Bounding Box Regression,” AAAI, 2020. DOI: 10.1609/aaai.v34i07.6999.
- G. Hinton, O. Vinyals, and J. Dean, “Distilling the Knowledge in a Neural Network,” arXiv:1503.02531, 2015.
- Y. Zhang et al., “ByteTrack: Multi-Object Tracking by Associating Every Detection Box,” ECCV, 2022.
Primary local evidence
docs/technical_report.mddocs/v7_audit_status.mdnotes.mdartifacts/v7_gates/official_test_confirmation/epoch33-paired-official-once-v3/report.jsonartifacts/cpu_student/hybrid_distilled_student_eval_latency_summary.jsonartifacts/cpu_student/hybrid_distilled_student_full_cpu_latency.jsonartifacts/cpu_student/post_training_latency_neutral/promotion.jsonartifacts/cpu_student/post_training_latency_neutral/tracker_live_accuracy.jsonartifacts/cpu_student/post_training_latency_neutral/tracker_live_latency.json







