|
Download docs/OrbitSight_Full_Technical_Proposal_Review.md from simerai/orbitsight: direct link, hf CLI and curl.
- Browser
- Download file 53.5 kB
-
https://huggingface.co/simerai/orbitsight/resolve/main/docs/OrbitSight_Full_Technical_Proposal_Review.md
- Command line
-
hf download hf://simerai/orbitsight/docs/OrbitSight_Full_Technical_Proposal_Review.md
-
curl -L -o OrbitSight_Full_Technical_Proposal_Review.md https://huggingface.co/simerai/orbitsight/resolve/main/docs/OrbitSight_Full_Technical_Proposal_Review.md
53.5 kB
| # OrbitSight — Full Technical Proposal and Engineering Review | |
| **Long-form review draft, 29 August 2026** | |
| **Team SIMER — Hishaam A., Simreen Siraj, and Christian Sabatini — Ampiere Labs** | |
| > This is deliberately not the five-page submission proposal. It is the detailed source | |
| > document from which a shorter competition proposal, pitch deck, or technical appendix can | |
| > be produced. It tells the complete development story, including failed models, invalid | |
| > experiments, latency failures, and the limits of the available evaluation evidence. | |
| ## Executive summary | |
| OrbitSight detects and tracks resident space objects (RSOs) directly from asynchronous | |
| neuromorphic vision sensor events. The final submission model is a compact recurrent neural | |
| network followed by a physically interpretable Kalman tracker. It reads raw `*.npy` event | |
| recordings, processes consecutive 40 ms windows, and writes the exact ChallengeON prediction | |
| files and evaluation workbook from an automatic, offline, CPU-only Docker container. | |
| The important story is not a single final score. It is a sequence of engineering discoveries: | |
| 1. The first honest pipeline scored only **F1 0.036 / mAP 0.010**. Early capped experiments | |
| had silently missed most labelled windows, and several apparently good measurements were | |
| not reproducible. | |
| 2. Correct windowing, event denoising, a three-channel representation, recurrent detection, | |
| CIoU localization, longer temporal training, and honest best-epoch selection produced V5. | |
| After output calibration and gate corrections, V5 reached approximately **F1 0.630 / mAP | |
| 0.380** on the leakage-free local diagnostic. | |
| 3. V7 introduced a more principled area-resize and coordinate-mapping contract. It did not | |
| beat V5 overall, but it was substantially better on EVK4. V5 remained better on DVX. | |
| 4. A sensor router combined the specialists: **EVK4 → V7**, **DAVIS/DVX → V5**. This hybrid | |
| reached **F1 0.668 / mAP 0.490** leakage-free and **0.676 / 0.549** across all four local | |
| recordings. The route was chosen after per-sensor results were visible, so it is a | |
| score-maximizing diagnostic rather than blind model-selection evidence. | |
| 5. The heavy hybrid was unusable on the competition CPU: approximately **383 ms mean / 451 ms | |
| p95**, around an order of magnitude above the 40 ms requirement. | |
| 6. The hybrid therefore became a teacher rather than a deployable model. A 135,111-parameter | |
| recurrent student was trained from human labels plus carefully filtered teacher evidence. | |
| It retained full 640×640 localization while being approximately **63.6× smaller** than one | |
| 8,588,339-parameter teacher branch. | |
| 7. The distilled student measured **F1 0.6229 / mAP 0.4608** leakage-free and **0.6375 / | |
| 0.5373** on all four recordings, with **18.55 ms mean / 21.97 ms worst-sequence p95** in a | |
| full local CPU profile. | |
| 8. Without retraining, conservative sensor calibration and tracker changes raised the final | |
| diagnostics to **F1 0.6292 / mAP 0.4645** leakage-free and **0.6437 / 0.5410** all-four. | |
| A later 3,920-window profile measured **17.05 ms mean / 20.64 ms worst-sequence p95**, with | |
| the largest observed window at 39.17 ms. | |
| The result is not the highest-scoring network we built. It is the strongest practical balance | |
| of detection quality, CPU latency, reproducibility, and challenge-format correctness. | |
| ## 1. Problem and evaluation objective | |
| ### 1.1 What the sensor produces | |
| A conventional camera exposes a dense image at fixed intervals. A neuromorphic vision sensor | |
| works differently: each pixel emits an event when local log-brightness changes enough. An | |
| event carries pixel position, polarity, and a microsecond timestamp. The result is a sparse, | |
| asynchronous stream rather than a sequence of ordinary photographs. | |
| This is attractive for space situational awareness because event sensors offer high dynamic | |
| range, low motion blur, and fine timing. It is also difficult because most events are not the | |
| target. In the supplied recordings, labelled RSO events represent roughly **0.08–0.96%** of | |
| the stream. The rest includes sensor activity, hot pixels, star-field responses, and other | |
| background structure. | |
| OrbitSight supports three sensor families with very different native resolutions: | |
| | Sensor | Native resolution | Transformation challenge | | |
| |---|---:|---| | |
| | DAVIS346c | 346×260 | Must be enlarged while preserving tiny target geometry | | |
| | DVXplorer | 640×480 | Near the model scale, but strongly affected by low-SNR clutter | | |
| | EVK4 | 1280×720 | Must be reduced without deleting sparse events | | |
| ### 1.2 Why the targets are unusually difficult | |
| - The median labelled box is only about **10×11 pixels**; the tenth percentile is about | |
| **6×6 pixels**. | |
| - IoU ≥ 0.5 on a ten-pixel object allows only a few pixels of center error. | |
| - Signal strength varies by roughly 200× across recordings. | |
| - Every annotated window contains one target box, so a second emitted detection cannot become | |
| another true positive; it can only reduce precision. | |
| - The hardest low-SNR cases contain only about six useful object events in a 40 ms window. | |
| At that point, localization is limited by sampling noise as much as by model capacity. | |
| ### 1.3 What must be delivered | |
| For every input sequence, the submission produces `<sequencename>.txt` with these fields: | |
| `sequence_id, window_start_timestamp_us, window_end_timestamp_us, x_centre, y_centre, w, h, class_id, confidence` | |
| The container also produces `Evaluation_Metrics.xlsx`. Accuracy is reported using precision, | |
| recall, F1, and AP/mAP at IoU 0.5. The competition additionally calls for end-to-end latency | |
| below 40 ms on an Intel Core i9-12900H CPU. There is no track-ID column, so tracking matters | |
| only when it improves the emitted detection stream. | |
| ## Literature review and design lineage | |
| OrbitSight is an application-specific synthesis rather than a direct reproduction of one published | |
| network. Its design draws on neuromorphic sensing, recurrent event detection, anchor-free dense | |
| detection, model compression, and detection-assisted tracking. | |
| ### Event cameras and space situational awareness | |
| Gallego et al. formalize the event-camera model: a pixel reports a change in log intensity rather | |
| than an absolute image sample. The literature emphasizes microsecond timing, high dynamic range, | |
| low motion blur, and sparse output, while warning that frame algorithms require an event | |
| representation or event-native computation. OrbitSight preserves polarity and within-window timing, | |
| but packs events into 40 ms tensors so convolutional and recurrent operations run efficiently on a | |
| CPU. [Gallego et al., *Event-based Vision: A Survey*, TPAMI 2020](https://doi.org/10.1109/TPAMI.2020.3008413). | |
| Afshar et al. demonstrated event-based RSO detection and tracking across sensors and observing sites. | |
| Ralph et al. subsequently studied astrometric calibration and source characterization for event-based | |
| space imaging. These works motivate OrbitSight's explicit sensor profiles, native-coordinate | |
| reconstruction, and treatment of preprocessing and coordinate mapping as part of the model contract. | |
| [Afshar et al., *Event-based Object Detection and Tracking for Space Situational Awareness*, | |
| 2019](https://arxiv.org/abs/1911.08730); [Ralph et al., *Astrometric Calibration and Source | |
| Characterisation of Neuromorphic Event-based Cameras for Space Imaging*, | |
| 2022](https://arxiv.org/abs/2211.09939). | |
| ### Recurrent event-based object detection | |
| Perot et al.'s Recurrent Event-camera Detector showed that direct event representations combined with | |
| recurrent memory outperform feed-forward event detectors because state integrates evidence through | |
| time. RVT later showed that multi-stage recurrent backbones can obtain a favorable latency/accuracy | |
| tradeoff. OrbitSight adopts direct event tensors, multi-scale features, and persistent temporal state, | |
| but uses narrow convolutions and ConvLSTM instead of a transformer for an offline CPU target and tiny | |
| objects. [Perot et al., *Learning to Detect Objects with a 1 Megapixel Event Camera*, NeurIPS | |
| 2020](https://proceedings.neurips.cc/paper/2020/hash/c213877427b46fa96cff6c39e837ccee-Abstract.html); | |
| [Gehrig and Scaramuzza, *Recurrent Vision Transformers for Object Detection With Event Cameras*, | |
| CVPR 2023](https://openaccess.thecvf.com/content/CVPR2023/html/Gehrig_Recurrent_Vision_Transformers_for_Object_Detection_With_Event_Cameras_CVPR_2023_paper.html). | |
| ConvLSTM replaces fully connected state transitions with spatial convolutions, retaining *where* weak | |
| evidence occurred. OrbitSight places one after every downsampling stage and uses dilations 1, 2, 4, | |
| and 8 to enlarge spatial context without changing tensor shapes. [Shi et al., *Convolutional LSTM | |
| Network*, NeurIPS 2015](https://proceedings.neurips.cc/paper/2015/hash/07563a3fe3bbe7e3ba84431ad9d055af-Abstract.html). | |
| ### Anchor-free detection, imbalance, and localization | |
| FCOS established fully convolutional anchor-free prediction; YOLOX paired anchor-free boxes with a | |
| decoupled head. OrbitSight adopts those ideas with custom assignment and scoring for one RSO class. | |
| It predicts raw center offsets and log sizes on three scales without an anchor bank. [Tian et al., | |
| *FCOS*, ICCV 2019](https://openaccess.thecvf.com/content_ICCV_2019/html/Tian_FCOS_Fully_Convolutional_One-Stage_Object_Detection_ICCV_2019_paper.html); | |
| [Ge et al., *YOLOX*, 2021](https://arxiv.org/abs/2107.08430). | |
| Dense event tensors create tens of thousands of background cells for at most one target. Focal loss | |
| suppresses easy negatives so hard examples control learning. OrbitSight uses sigmoid focal BCE for | |
| objectness. CIoU combines overlap, center distance, and aspect-ratio consistency, aligning training | |
| more closely with the IoU≥0.5 metric than independent coordinate L1. [Lin et al., *Focal Loss for | |
| Dense Object Detection*, ICCV 2017](https://openaccess.thecvf.com/content_ICCV_2017/papers/Lin_Focal_Loss_for_ICCV_2017_paper.pdf); | |
| [Zheng et al., *Distance-IoU Loss*, AAAI 2020](https://ojs.aaai.org/index.php/AAAI/article/view/6999). | |
| ### Distillation and tracking | |
| Knowledge distillation transfers useful behavior from an expensive model or ensemble into a smaller | |
| deployable network. OrbitSight follows that deployment logic but does not blindly copy teacher output: | |
| human boxes remain authoritative, matched teacher boxes provide low-weight targets, and proposals in | |
| human-empty windows become negative evidence. [Hinton, Vinyals, and Dean, *Distilling the Knowledge | |
| in a Neural Network*, 2015](https://arxiv.org/abs/1503.02531). | |
| ByteTrack showed the value of associating lower-confidence detections. OrbitSight evaluated that | |
| two-tier idea, but production emits only stable global top-one at 0.10, making the nominal low tier | |
| unreachable. The final system is therefore a single-tier constant-velocity Kalman/coast tracker using | |
| the same association machinery. [Zhang et al., *ByteTrack*, ECCV | |
| 2022](https://www.ecva.net/papers/eccv_2022/papers_ECCV/papers/136820001.pdf). | |
| ### OrbitSight's design contribution | |
| - a polarity/recency representation with sensor-specific event denoising; | |
| - a full-resolution but ultra-narrow recurrent detector; | |
| - single-class objectness decoding without a redundant class multiplier; | |
| - selective, human-authoritative distillation from a sensor-routed teacher; | |
| - weights bound to preprocessing and coordinate-mapping fingerprints; | |
| - exact stable top-one decoding that avoids NMS without changing the selected box; | |
| - one evaluated detector/tracker/writer deployment contract. | |
| ## 2. Development history at a glance | |
| The following table separates comparable leakage-free results from historical development | |
| measurements. The `detector_e30` row came from the validation split used at that time and is | |
| therefore useful as progress evidence, not a strict comparison with later test rows. | |
| | Stage | Main change | Leakage-free F1 | Leakage-free mAP | Decision | | |
| |---|---|---:|---:|---| | |
| | First honest baseline | Correctly measured initial pipeline | 0.036 | 0.010 | Foundation only | | |
| | `detector_e30` | 30-epoch trained pipeline and tuned tracking | 0.515* | 0.220* | Promising, not directly comparable | | |
| | `detector_ciou` | CIoU, dilations, regime balancing | 0.5620 | 0.2536 | Retained ideas | | |
| | V5 detector | TBPTT-40, honest split, best epoch | 0.5808 | 0.2749 | Strong detector checkpoint | | |
| | V5 + box calibration | Per-sensor size correction | 0.5933 | 0.2947 | Retained | | |
| | V5 complete pipeline | Low detector gate and per-sensor emit gates | ≈0.6300 | ≈0.3798 | Scientific baseline | | |
| | V6 | Intended area-resize successor | 0.6116† | 0.3570† | Quarantined: invalid contract | | |
| | V7 epoch 33 | Correct area resize and pixel-center mapping | 0.5954 | 0.3585 | Lost overall; EVK4 specialist | | |
| | Routed V5/V7 hybrid | EVK4 V7, DAVIS/DVX V5 | 0.6683 | 0.4895 | Accuracy teacher; too slow | | |
| | Compact student from scratch | Same small architecture, no teacher | 0.1055 aggregate | 0.2602 aggregate | DVX collapsed | | |
| | Distilled CPU student | Human labels + selective hybrid teaching | 0.6229 | 0.4608 | Fast deployable candidate | | |
| | Final CPU student | No-retrain calibration and tracker refinement | **0.6292** | **0.4645** | Recommended submission | | |
| \* Historical validation regime. | |
| † Retained only as an audit trail because training inputs and declared preprocessing did not match. | |
|  | |
| Two lessons are visible immediately. First, the large gain from below 0.1 to approximately | |
| 0.63 did not come from one architectural trick; it came from correcting the entire data, | |
| training, inference, tracking, and writing contract. Second, the best accuracy model was not | |
| the best deployment model. The heavy hybrid forced a separate compression phase. | |
| ## 3. Why the first scores were below 0.1 | |
| ### 3.1 The first honest result | |
| The first reproducible leakage-free baseline was **F1 0.0364 / mAP 0.0096**. This was useful | |
| because it established an honest floor, even though the score was extremely low. An earlier | |
| F1 value around 0.14 came from a measurement path that could not be reproduced consistently | |
| and is not used as the project baseline. | |
| ### 3.2 The label-starvation bug | |
| Early capped experiments evaluated the first N windows of a stream. Labels, however, often | |
| began much later. In one sequence the first labelled window was 1,562, so a 1,500-window cap | |
| could finish before reaching the main annotated region. Across the affected capped runs, | |
| only about 34% of labels were visible. | |
| This bug mattered twice: training batches contained too few positive examples, and evaluation | |
| looked at the wrong temporal region. The fix introduced an explicit starting window and | |
| aligned capped runs to the first labelled window. This was not a model improvement; it was a | |
| measurement correction that made later model improvements meaningful. | |
| ### 3.3 Other foundational corrections | |
| - Ground-truth timestamps were required to lie on the exact 40 ms window phase. | |
| - The official evaluator's pixel-inclusive IoU convention was reproduced locally. | |
| - Prediction box fields were made integer-valued only at the final writer boundary. | |
| - Zero-sized boxes after rounding were clamped to at least one pixel. | |
| - Model objectness was decoded correctly for the single-class RSO task instead of being | |
| multiplied by a redundant class probability. | |
| - Detector, recurrent, and tracker state were reset at every sequence boundary. | |
| These corrections are easy to underestimate. A sophisticated model evaluated through a | |
| misaligned or inconsistent pipeline is still a bad system. | |
| ## 4. The common event-processing pipeline | |
| The heavy models and compact student share the same conceptual flow: | |
| ```text | |
| raw event stream | |
| ↓ | |
| memory-mapped ingestion and phase-checked 40 ms windows | |
| ↓ | |
| sensor-specific background-activity filtering | |
| ↓ | |
| three-channel event tensor at 640×640 | |
| ↓ | |
| recurrent anchor-free detector | |
| ↓ | |
| map box back to native sensor coordinates | |
| ↓ | |
| Kalman association and bounded coasting | |
| ↓ | |
| sensor-specific output gate and one-box cap | |
| ↓ | |
| integer portal row + Evaluation_Metrics.xlsx | |
| ``` | |
| ### 4.1 Background-activity filtering | |
| A real optical signal tends to produce neighboring events close together in time. Independent | |
| sensor noise is less likely to have that support. The BA filter asks a simple question for | |
| each event: did enough nearby pixels fire recently? The parameters are sensor-specific: | |
| | Sensor | Required support | Radius | Time horizon | RSO events retained | Noise rejected | | |
| |---|---:|---:|---:|---:|---:| | |
| | DAVIS | 1 | 6 px | 40 ms | 99.47% | 28.32% | | |
| | DVX | 1 | 4 px | 20 ms | 98.24% | 39.94% | | |
| | EVK4 | 1 | 3 px | 5 ms | 97.71% | 41.41% | | |
| The filter is window-local and stateless in the deployed route. Hot-pixel support exists but | |
| is not silently enabled because changing preprocessing only at inference would create a | |
| train/deployment mismatch. | |
| ### 4.2 Three-channel representation | |
| Each 40 ms event window becomes a three-channel tensor: | |
| 1. positive-event count; | |
| 2. negative-event count; | |
| 3. recency of the latest event. | |
| The first two channels show where brightness increased or decreased. The recency channel | |
| preserves temporal direction inside the window. A moving point source often produces a | |
| polarity pattern and timestamp gradient, so these three channels retain more useful physics | |
| than a single accumulated image. | |
| The model input remains 640×640 even in the compact student. Reducing the image would save | |
| compute, but a 6–11 pixel target cannot tolerate much spatial quantization. We reduced network | |
| width instead of discarding localization resolution. | |
| ## 5. From early trained models to V5 | |
| ### 5.1 The detector family | |
| OrbitSight uses an anchor-free, multi-scale detector. “Anchor-free” means that the network | |
| predicts a target center and size directly at grid locations instead of choosing from a bank | |
| of predefined rectangle shapes. This suits small RSOs whose apparent extent depends on | |
| brightness, motion, optics, and sensor resolution. | |
| The backbone contains four downsampling stages. Each stage includes ConvLSTM recurrence, so | |
| the model carries information from earlier windows. Recurrent dilations `(1, 2, 4, 8)` expand | |
| the spatial context without requiring a much deeper network. Detection heads operate at | |
| multiple scales, with stride 4 retained for precise small-target localization. | |
| ### 5.2 CIoU and confidence learning | |
| Early box regression did not optimize the competition's overlap criterion closely enough. | |
| Complete IoU (CIoU) improved the learning signal by combining overlap, center distance, and | |
| shape consistency. It also supplies a gradient when a predicted box and target do not yet | |
| overlap. | |
| The objectness target is linked to achieved localization quality. In plain language: a box | |
| should be confident when it is both likely to contain the RSO and accurately localized. This | |
| makes confidence useful for ranking, but it also means correct tiny boxes may have moderate | |
| rather than near-one scores. That observation later motivated separating the detector's low | |
| proposal threshold from the final output gate. | |
| ### 5.3 Training improvements that produced V5 | |
| - **Regime balancing:** batches were designed to prevent abundant background windows from | |
| erasing rare positive examples. | |
| - **TBPTT length 40:** the network learned across 1.6 seconds of 40 ms windows instead of a | |
| shorter temporal horizon. | |
| - **Honest validation:** sequence-level separation reduced leakage from adjacent windows. | |
| - **Best-epoch selection:** V5 kept epoch 33 rather than simply shipping the final epoch. | |
| - **Full-path validation:** evaluation included tracking, coasting, gates, and writer | |
| quantization—the path that would actually be submitted. | |
| V5's heavy backbone widths are `(32, 64, 128, 256)`, the head width is 128, and the model has | |
| **8,588,339 parameters**. It became the scientific baseline because it was the strongest | |
| model supported by a coherent, reproducible contract. | |
| ### 5.4 The output-gate breakthrough | |
| Originally, the detector threshold and final output threshold were effectively the same. A | |
| box below the emit threshold was discarded before tracking, even if it was the best proposal | |
| in the window. The corrected design runs the detector at a low 0.10 proposal threshold, lets | |
| the tracker use that evidence, and applies the sensor output gate only at the writer. | |
| This separation was one of the largest no-retrain gains in the heavy pipeline. The retained | |
| V5 gates were DAVIS 0.30, DVX 0.10, and EVK4 0.10. Per-sensor box-size correction also repaired | |
| systematic under-sizing caused by sensor transformations. Together, these changes moved the | |
| complete V5 pipeline to approximately **F1 0.630 / mAP 0.380** leakage-free. | |
| ## 6. V6 and the importance of a preprocessing contract | |
| V6 was intended to test area resampling and add dim-regime data. Its recorded score appeared | |
| competitive, but an audit found a critical mismatch: training consumed historical nearest- | |
| resized tensors while metadata claimed area resize. Evaluation could therefore feed one | |
| representation to weights trained on another. | |
| The checkpoint still loaded because tensor shapes matched. That is exactly why the problem | |
| was dangerous: shape compatibility did not imply semantic compatibility. V6 was quarantined, | |
| and later loaders were strengthened to bind checkpoints to preprocessing fingerprints. | |
| The general lesson is central to OrbitSight: a model is not just weights. It is weights plus | |
| window definition, denoising, resize method, coordinate mapping, count transform, state | |
| behavior, thresholds, and writer projection. | |
| ## 7. V7: better geometry, mixed sensor behavior | |
| V7 rebuilt the preprocessing contract around true area resampling and pixel-center rounding. | |
| The intent was sound: when EVK4 is reduced from 1280×720, nearest-neighbor sampling can delete | |
| sparse events, while area resampling preserves aggregate activity more faithfully. | |
| The consolidated V7 model did not beat V5 overall: | |
| | Model | Leakage-free F1 | Leakage-free mAP | All-four F1 | All-four mAP | | |
| |---|---:|---:|---:|---:| | |
| | Paired V5 | 0.6298 | 0.3804 | 0.6403 | 0.4670 | | |
| | V7 epoch 33 | 0.5954 | 0.3585 | 0.6100 | 0.4378 | | |
| However, aggregate results hid meaningful specialization: | |
| | Sequence | V5 F1 / AP | V7 F1 / AP | Better branch | | |
| |---|---:|---:|---| | |
| | EVK4 magnitude 7.3 | 0.5175 / 0.2797 | 0.7121 / 0.6070 | V7 by a large margin | | |
| | DVX Stars3 | 0.6667 / 0.5812 | 0.5723 / 0.3525 | V5 | | |
| | DVX Thuraya3 | 0.4357 / 0.2803 | 0.2095 / 0.1160 | V5 | | |
| V7 was therefore not a universal successor. It was an EVK4 specialist. This distinction | |
| created the hybrid teacher. | |
| ## 8. The V5/V7 sensor-routed hybrid | |
| The router chooses one heavy branch once per sequence: | |
| - EVK4 → V7 epoch 33; | |
| - DAVIS → V5 best; | |
| - DVX → V5 best. | |
| Only one branch runs for a given window, so this is routing rather than simultaneous | |
| ensembling. The routing result was: | |
| | Scope | Precision | Recall | F1 | mAP | | |
| |---|---:|---:|---:|---:| | |
| | Leakage-free | 0.6400 | 0.6991 | 0.6683 | 0.4895 | | |
| | All four | 0.6504 | 0.7036 | 0.6760 | 0.5489 | | |
| These were the highest practical local accuracy values in the project. They are also | |
| post-hoc: the route was proposed after the per-sensor results were known. The hybrid is | |
| therefore useful as a score-maximizing package and as a teacher, but its local result is not | |
| an unbiased estimate of unseen performance. | |
| ### 8.1 Why the hybrid could not be submitted on CPU | |
| Each branch still has 8.59 million parameters and full-width recurrent computation. On a | |
| local four-thread CPU, the hybrid measured approximately: | |
| - **383 ms mean per window**; | |
| - **451 ms p95 per window**; | |
| - roughly **313 ms/window effective** during the complete four-sequence replay. | |
| The budget is 40 ms. This is not a small optimization gap. Tracker tuning, output formatting, | |
| or minor Python changes cannot create a tenfold speedup in the dominant convolutional model. | |
| The architecture had to become much smaller. | |
| ## 9. Why teacher–student distillation was chosen | |
| The hybrid knew useful sensor-specific behavior, but was too slow to deploy. Distillation | |
| separates those roles: | |
| - the **teacher** is allowed to be large and expensive because it runs only while generating | |
| training evidence; | |
| - the **student** learns from human labels plus selected teacher information, then runs alone | |
| in the submission container. | |
| The shipped image does not contain V5, V7, or the router. Teacher latency is therefore absent | |
| from deployment latency. | |
| ### 9.1 Why a small model trained from scratch was not enough | |
| The same compact architecture was first trained without the routed teacher. Its aggregate | |
| F1/mAP was **0.1055 / 0.2602** and DVX effectively collapsed. On a six-positive DVX held-out | |
| capture, the best examined epoch produced 2 TP, 3,941 FP, and 4 FN. The two true positives | |
| ranked 793rd and 2,595th by confidence. | |
| No output threshold can repair that ordering. Raising the threshold removes the buried true | |
| positives along with false positives; lowering it emits thousands of false detections. The | |
| failure was a hard-negative ranking problem, which is exactly where a stronger teacher can | |
| provide useful structure. | |
| ### 9.2 Selective distillation rather than blind imitation | |
| Teacher outputs are not automatically true. The training design keeps human ground truth | |
| authoritative: | |
| - A teacher detection that matches human GT may contribute a low-weight soft distillation | |
| target. | |
| - A human GT missed by the teacher remains a full positive. | |
| - A teacher detection in a human-empty window is treated as hard-negative/ranking evidence, | |
| not copied as a positive. | |
| - Post-tracker teacher boxes are diagnostics only; the student learns pre-tracker detection | |
| behavior so it is not tracked twice. | |
| - Teacher recurrence, student recurrence, BA state, and tracking state reset at sequence | |
| boundaries. | |
| - Cache manifests bind source hashes, sensor parameters, preprocessing, coordinate mapping, | |
| checkpoint identity, and code/schema versions. | |
| This design prevents the most dangerous distillation failure: teaching the student to repeat | |
| the teacher's false positives. | |
| ### 9.3 Training evidence and run | |
| Teacher evidence was generated in two independent shards and merged under signed manifests. | |
| Across the merged cache there were: | |
| - 17 training sequences; | |
| - 106,194 windows; | |
| - 18,139 pre-tracker teacher detections; | |
| - 16,682 post-tracker diagnostic detections; | |
| - 11,343 positive windows with a matching teacher detection; | |
| - 5,238 human-empty windows containing teacher proposals, retained as hard negatives. | |
| The final A100 training used 36 epochs, 260 updates per epoch, 9,360 optimizer updates, batch | |
| 16, and segment length 20. An interrupted client connection was recovered by exact checkpoint | |
| resume, including model, optimizer, scheduler, RNG, sampler, and batching state. The final | |
| checkpoint SHA-256 is: | |
| `5f176b9baffc3478b1aaada13277e2c6093ba02439b76b8330cb8d654f5f554d` | |
| ## 10. Final-model methodology | |
| This section describes how the final student was produced and evaluated. Section 11 then defines | |
| the frozen network and inference path component by component. | |
| ### 10.1 Research questions and temporal unit | |
| Development asked whether one model could localize tiny RSOs across all three sensor families, | |
| inherit useful sensor-specific behavior from the heavy hybrid, and keep the complete deployed path | |
| below 40 ms. The atomic example is a half-open 40,000 µs event interval. Window phase is anchored to | |
| the sequence timestamp contract. Recurrent, BA, and tracker state are initialized at the start of a | |
| sequence, advanced exactly once per window, and destroyed at the boundary. | |
| Training uses ordered segments, not shuffled independent frames. Final distillation used 20 consecutive | |
| windows per segment—0.8 seconds of temporal context—and batches of 16 independent segments. Each batch | |
| member carries its own four hidden/cell state pairs before truncated backpropagation detaches them. | |
| ### 10.2 Data preparation | |
| For every window the method: | |
| 1. memory-maps and timestamp-slices events; | |
| 2. applies the sensor-specific neighborhood/time BA rule; | |
| 3. builds positive-count, negative-count, and latest-event-recency planes; | |
| 4. area-resizes the planes to 640×640; | |
| 5. maps human boxes with the same pixel-center convention; | |
| 6. builds the optional per-event auxiliary target. | |
| Counts remain raw. Stateful BA, sensor FiLM, hot-pixel changes, and target-event thinning are disabled | |
| in the final route. Cache fingerprints bind the source, sensor profile, 40 ms phase, input size, count | |
| transform, area resize, and coordinate mapping. This prevents a shape-compatible checkpoint from being | |
| used with semantically different preprocessing, the failure that invalidated V6. | |
| ### 10.3 Human-authoritative teacher evidence | |
| The V5/V7 router runs offline. Its pre-tracker proposals are stored in sidecars aligned one-to-one with | |
| student windows and marked as matched or unmatched against human GT. Post-tracker teacher boxes are | |
| diagnostic only. The selective objective is: | |
| ```text | |
| L_total = L_human | |
| + 0.15 L_teacher-objectness-on-matched-GT | |
| + 0.25 L_teacher-box-on-matched-GT | |
| + 0.35 L_teacher-empty-window-negative | |
| ``` | |
| Teacher scores below 0.05 are ignored. Matched proposals supply a BCE confidence target and Smooth-L1 | |
| encoded-box target. A proposal in a human-empty window pushes the corresponding student objectness | |
| toward zero. An unmatched teacher proposal never creates a positive label. The teacher is therefore a | |
| ranking guide and hard-negative miner, not a replacement annotator. | |
| ### 10.4 Optimization protocol | |
| The final A100 run used 36 epochs, 260 updates per epoch, and 9,360 optimizer steps: | |
| | Item | Final setting | | |
| |---|---:| | |
| | Optimizer | Adam | | |
| | Initial / minimum LR | 5×10⁻⁴ / 1×10⁻⁵ | | |
| | Warm-up / schedule | 100 steps / cosine decay | | |
| | Precision | FP32; TF32 disabled | | |
| | Gradient-norm clip | 1.0 | | |
| | Batch / segment | 16 / 20 windows | | |
| | Negative/positive target | 2.0 | | |
| | Maximum oversampling | 4.0× | | |
| | Human box loss | pixel-inclusive CIoU, weight 2.0 | | |
| | Auxiliary segmentation | BCE, weight 0.5 | | |
| | Objectness target floor | 1.0 throughout | | |
| | Quality head / loss | disabled / 0.0 | | |
| | Seed | 1337 | | |
| Adam state, scheduler, RNG, batching order, partial epoch progress, preprocessing fingerprint, and the | |
| training contract were included in exact-resume checkpoints. A remote disconnect therefore resumed | |
| the same stochastic run rather than launching a similar but different experiment. | |
| ### 10.5 Evaluation methodology | |
| Predictions are scored only after native mapping, calibration, tracking, clipping, top-one selection, | |
| sensor output gating, and integer writer projection. IoU uses the official pixel-inclusive convention. | |
| Metrics are reported for the three-sequence leakage-free aggregate and all four local recordings. | |
| Because every remaining recording has influenced diagnostics, these are reproducibility and local | |
| comparison results—not blind-generalization proof. Latency is measured after warm-up at streaming | |
| batch one and reports mean, per-sequence p95, and maximum rather than forward-pass mean alone. | |
| ## 11. Final model architecture: complete component reference | |
| ### 11.1 Frozen architecture contract | |
| | Property | Final value | | |
| |---|---:| | |
| | Class | `OrbitSightDetector` | | |
| | Inference input | 1×3×640×640 | | |
| | Backbone widths | 4, 8, 16, 32 | | |
| | ConvLSTM dilations | 1, 2, 4, 8 | | |
| | Detection scales | stride 4, 8, 16 | | |
| | Head width | 16 | | |
| | Classes | one RSO foreground class | | |
| | Stored parameters | 135,111 | | |
| | Quality head / sensor FiLM | absent / absent | | |
| | Auxiliary segmentation | stored; training-only execution | | |
| | Classification towers | stored for compatibility; never executed | | |
|  | |
| ### 11.2 Compression decision | |
| | Component | Heavy V5/V7 branch | CPU student | | |
| |---|---:|---:| | |
| | Input | 3×640×640 | 3×640×640 | | |
| | Backbone widths | 32, 64, 128, 256 | 4, 8, 16, 32 | | |
| | Head width | 128 | 16 | | |
| | Dilations | 1, 2, 4, 8 | 1, 2, 4, 8 | | |
| | Parameters | 8,588,339 | 135,111 | | |
| | Runtime role | Offline teacher | Deployed model | | |
|  | |
| The 63.6× reduction comes mainly from channel width. Input resolution, recurrence, the finest stride-4 | |
| grid, and three detection scales remain because spatial quantization is expensive for 6–11 pixel targets. | |
| ### 11.3 Input representation | |
| `X ∈ R^(B×3×640×640)` is float32. Channel 0 counts positive-polarity events; channel 1 counts negative | |
| events; channel 2 records normalized recency of the latest event. Empty pixels are zero. This preserves | |
| brightness-change sign and within-window order without reconstructing an intensity image. | |
| ### 11.4 Stem | |
| ```text | |
| (B,3,640,640) | |
| → Conv2d(3→4, kernel 3, stride 1, pad 1) | |
| → BatchNorm2d(4) → SiLU | |
| = (B,4,640,640) | |
| ``` | |
| The stem has 120 parameters. SiLU is `x·sigmoid(x)`. There is no pooling or tokenization. | |
| ### 11.5 Four downsampling and recurrent stages | |
| Each stage performs 3×3 stride-2 convolution, BatchNorm, SiLU, then ConvLSTM. There are no residual | |
| blocks, attention layers, transformer tokens, or active sensor-conditioning layers. | |
| | Stage | Down mapping | Feature and state | Dilation | Down params | ConvLSTM params | Head input? | | |
| |---|---|---|---:|---:|---:|---| | |
| | 1 | 4→4 | B×4×320×320 | 1 | 156 | 1,168 | No | | |
| | 2 | 4→8 | B×8×160×160 | 2 | 312 | 4,640 | Yes, stride 4 | | |
| | 3 | 8→16 | B×16×80×80 | 4 | 1,200 | 18,496 | Yes, stride 8 | | |
| | 4 | 16→32 | B×32×40×40 | 8 | 4,704 | 73,856 | Yes, stride 16 | | |
| Stage 1 still matters: its high-resolution recurrent output feeds all later stages. The last three | |
| tensors are independent multi-scale features. Unlike a classical FPN, there is **no top-down pathway, | |
| lateral addition, or cross-scale fusion**. | |
| ### 11.6 Every ConvLSTM operation | |
| At time `t`, current feature `X_t` and previous hidden state `H_(t−1)` are concatenated. One dilated | |
| 3×3 convolution emits four channel groups: | |
| ```text | |
| [I,F,O,G] = Conv3×3_dilated([X_t,H_(t−1)]) | |
| i=sigmoid(I), f=sigmoid(F), o=sigmoid(O), g=tanh(G) | |
| C_t = f ⊙ C_(t−1) + i ⊙ g | |
| H_t = o ⊙ tanh(C_t) | |
| ``` | |
| The input, forget, and output gates govern writing, retaining, and exposing memory. Both `H_t` and | |
| `C_t` persist. At inference the four pairs have shapes 1×4×320×320, 1×8×160×160, | |
| 1×16×80×80, and 1×32×40×40. Dilation gives effective 3×3, 5×5, 9×9, and 17×17 footprints without | |
| adding weights. State begins at zero and resets only at sequence boundaries. | |
| ### 11.7 Three active detection heads | |
| For each returned feature, the active regression/objectness tower is: | |
| ```text | |
| feature C×H×W | |
| → Conv2d(C→16,3×3,pad1) → SiLU | |
| → Conv2d(16→16,3×3,pad1) → SiLU | |
| ├→ Conv2d(16→4,1×1): raw box vector | |
| └→ Conv2d(16→1,1×1): objectness logit | |
| ``` | |
| | Scale | Feature | Tower + box + objectness params | Outputs | | |
| |---|---:|---:|---| | |
| | stride 4 | 8×160×160 | 3,488 + 68 + 17 | 4×160×160 and 1×160×160 | | |
| | stride 8 | 16×80×80 | 4,640 + 68 + 17 | 4×80×80 and 1×80×80 | | |
| | stride 16 | 32×40×40 | 6,944 + 68 + 17 | 4×40×40 and 1×40×40 | | |
| There are 25,600 + 6,400 + 1,600 = **33,600 candidate cells**. Four regression channels encode raw | |
| x/y offsets and log width/height. Objectness alone asks whether an RSO occupies the cell. | |
| ### 11.8 Compatibility classification modules | |
| Three classification towers and 1×1 predictors remain in the checkpoint so historical state-dictionary | |
| keys load strictly. They contain 15,123 parameters, but `compute_cls=False`: their convolutions are not | |
| executed, their logits are not trained by the final loss, and the decoder never reads them. In a | |
| single-class problem, multiplying objectness by a redundant class probability would square/compress | |
| confidence without adding information. Final score is `sigmoid(objectness)`. No quality head exists. | |
| ### 11.9 Auxiliary segmentation head | |
| A nine-parameter 1×1 convolution maps the eight-channel stride-4 feature to one 160×160 logit plane. | |
| Per-event labels supervise it with BCE at weight 0.5, helping the finest features retain target-local | |
| structure. It runs only while `model.training` is true; deployment skips it entirely. | |
| ### 11.10 Target assignment and hard loss | |
| A human box is assigned to the scale whose `4×stride` nominal size is closest to its maximum dimension. | |
| The rounded center and radius-1 neighborhood—up to 3×3 cells—are positive. At cell `(g_x,g_y)`: | |
| ```text | |
| t_x=c_x/s−g_x, t_y=c_y/s−g_y, | |
| t_w=w/s, t_h=h/s | |
| ``` | |
| Objectness uses focal BCE with alpha 0.25 and gamma 2.0, normalized by positive assigned cells and | |
| clamped to one on empty windows. The positive objectness floor was 1.0 during final distillation. | |
| Regression uses pixel-inclusive CIoU: `1−IoU`, normalized center distance, and aspect-ratio penalty. | |
| It therefore supplies gradient even before two tiny boxes overlap. Auxiliary BCE and the selective KD | |
| terms described in Section 10 complete the objective. | |
| ### 11.11 Anchor-free decode | |
| For raw `(r_x,r_y,r_w,r_h)` at grid cell `(g_x,g_y)` and stride `s`: | |
| ```text | |
| c_x=(g_x+r_x)s c_y=(g_y+r_y)s | |
| w=exp(clamp(r_w,6))s h=exp(clamp(r_h,6))s | |
| score=sigmoid(objectness) | |
| ``` | |
| Offsets are intentionally raw, matching training. Exponentiation guarantees positive sizes and the | |
| upper clamp prevents overflow. | |
| ### 11.12 Exact stable top-one decode | |
| Production requests one proposal. The decoder concatenates three score maps in fixed scale/grid order, | |
| masks values below 0.10, and applies `argmax`. Exact ties select the first index deterministically. Only | |
| the winning cell's box is decoded. | |
| This exactly matches the first result of score-ordered NMS: the global maximum is selected before any | |
| overlap comparison can suppress it. Thus production avoids materializing/sorting 33,600 boxes and the | |
| Python NMS loop without approximating or changing the selected prediction. Generic pixel-inclusive NMS | |
| and its 1,000-candidate pre-bound remain only for multi-detection diagnostics. | |
| ### 11.13 Native-coordinate reconstruction | |
| The winning box is inverted through recorded area-resize geometry with pixel-center rounding. DVX | |
| width/height receives the frozen 0.98 calibration before association. Geometry is clipped in native | |
| sensor coordinates, but remains floating point until the writer performs final legal integer projection. | |
| ### 11.14 Complete parameter inventory | |
| | Module group | Parameters | Deployment execution | | |
| |---|---:|---| | |
| | Stem | 120 | Yes | | |
| | Four downsampling blocks | 6,372 | Yes | | |
| | Four ConvLSTM cells | 98,160 | Yes | | |
| | Three regression towers | 15,072 | Yes | | |
| | Box predictors | 204 | Yes | | |
| | Objectness predictors | 51 | Yes | | |
| | Compatibility classification modules | 15,123 | No | | |
| | Auxiliary segmentation | 9 | No | | |
| | **Total** | **135,111** | — | | |
| ### 11.15 End-to-end state machine | |
| For each sequence: identify sensor and reset state; slice 40 ms; denoise; form the 3×640×640 tensor; | |
| run stem and four ConvLSTMs; run three active heads; choose/decode global top-one; map/calibrate native | |
| geometry; update Kalman/coast state; retain at most one box; apply the sensor gate; project integer | |
| fields; write the row; and reset everything before another sequence. Docker latency covers this path, | |
| not merely the neural forward call. | |
| ## 12. Detection to tracking | |
| The detector proposes appearance-based boxes. The tracker imposes short-horizon physical | |
| consistency. Each track maintains: | |
| `[centre_x, centre_y, width, height, velocity_x, velocity_y]` | |
| Before the next detection arrives, a constant-velocity Kalman model predicts the new center | |
| and increases its uncertainty. If a detector box overlaps sufficiently, the state is updated. | |
| If evidence disappears briefly, the track may coast for a bounded number of windows with | |
| discounted confidence. | |
| The final student policy uses: | |
| - detector/high/low threshold 0.10; | |
| - one proposal before tracking and one output after tracking; | |
| - IoU association threshold 0.40; | |
| - `min_hits=1`, `max_age=15` internal lifetime; | |
| - at most one emitted coasted window; | |
| - coast confidence multiplied by 0.40; | |
| - no track IDs in the challenge file. | |
| This is intentionally simple. More elaborate motion gates were tested, but measured object | |
| motion was slow enough that consecutive boxes overlap. IoU association was therefore more | |
| reliable than a complex motion-only discriminator. | |
| ## 13. Distilled student results | |
| Before the final no-retrain adjustments, the distilled checkpoint produced: | |
| | Scope | Precision | Recall | F1 | mAP | TP | FP | GT | | |
| |---|---:|---:|---:|---:|---:|---:|---:| | |
| | Leakage-free | 0.5722 | 0.6834 | 0.6229 | 0.4608 | 4,472 | 3,343 | 6,544 | | |
| | All four | 0.5924 | 0.6899 | 0.6375 | 0.5373 | 4,882 | 3,359 | 7,076 | | |
| The student nearly matched V5 F1 while substantially improving the local mAP diagnostic, | |
| despite being 63.6× smaller than one heavy branch. It did not reach the aspirational F1 ≥ | |
| 0.70 and mAP > 0.60 targets; those targets also exceed the leakage-free hybrid teacher's | |
| measured values. | |
| The full 37,478-window CPU profile measured: | |
| - weighted mean 18.55 ms/window; | |
| - worst-sequence p95 21.97 ms/window; | |
| - every sequence p95 below 30 ms and 40 ms; | |
| - rare maximum outlier 152.66 ms. | |
| The last number is why both percentiles and maxima are reported. The system comfortably met | |
| the steady-state p95 objective locally, but scheduling outliers still existed in the longer | |
| profile. | |
| ## 14. Final no-retrain improvements | |
| With the checkpoint and architecture frozen, only low-cost constants were screened: | |
| | Setting | Previous student | Final student | | |
| |---|---:|---:| | |
| | DVX box-size factor | 1.00 | 0.98 | | |
| | DAVIS emit gate | 0.30 | 0.30 | | |
| | DVX emit gate | 0.10 | 0.10 | | |
| | EVK4 emit gate | 0.10 | 0.24 | | |
| | Tracker IoU association | 0.30 | 0.40 | | |
| | Maximum emitted coast | 2 windows | 1 window | | |
| | Coast confidence decay | 0.50 | 0.40 | | |
| Final measured accuracy: | |
| | Scope | Precision | Recall | F1 | mAP | Change in F1 | Change in mAP | | |
| |---|---:|---:|---:|---:|---:|---:| | |
| | Leakage-free | 0.5796 | 0.6881 | **0.6292** | **0.4645** | +0.0063 | +0.0037 | | |
| | All four | 0.5998 | 0.6946 | **0.6437** | **0.5410** | +0.0063 | +0.0037 | | |
|  | |
| The later 3,920-window profile measured 17.05 ms weighted mean, 20.64 ms worst-sequence | |
| p95, and 39.17 ms worst observed window. This shorter profile is not an apples-to-apples | |
| latency improvement claim against the earlier 37,478-window run; it confirms that the final | |
| policy remained comfortably under the local p95 limit and introduced no expensive operation. | |
| The gains are modest but consistent across both aggregate views. No weight, layer, or tensor | |
| shape changed. | |
| ## 15. Current per-sequence behavior | |
| | Recording | Sensor | F1 | AP | Precision | Recall | Main interpretation | | |
| |---|---|---:|---:|---:|---:|---| | |
| | EVK4 magnitude 7.3 | EVK4 | 0.7647 | 0.5927 | 0.7301 | 0.8027 | Strong benefit from EVK4-aware training behavior | | |
| | DAVIS SAOCOM-1B* | DAVIS | 0.8610 | 0.7708 | 0.9694 | 0.7744 | Strong but session-overlapping | | |
| | DVX Stars3 | DVX | 0.6014 | 0.4414 | 0.5419 | 0.6755 | Dominant false-positive source | | |
| | DVX Thuraya3 | DVX | 0.5166 | 0.3593 | 0.7569 | 0.3921 | Dimmest, recall-limited regime | | |
| \* SAOCOM-1B shares session context with training data and is excluded from the leakage-free | |
| aggregate. | |
| DVX remains the main unresolved accuracy problem. Stars3 contributes most false positives, | |
| while Thuraya3 has too little signal for adequate recall. These are not equivalent failures: | |
| one is primarily ranking/clutter, and the other is primarily missing evidence. | |
|  | |
|  | |
| ### 15.1 Animated detector-and-tracker demonstrations | |
| Four real-data animations accompany this review. Red is the raw top-one detector proposal, | |
| cyan is the box emitted after Kalman smoothing, and dashed yellow is human ground truth. The | |
| red and cyan trails show recent detector and tracker centers. Each frame also reports confidence, | |
| detector IoU, tracker IoU, event count, track ID/age, and elapsed time. | |
| - [EVK4 bright-target animation](../competition_ready_improved_cpu_student/visualizations/animations/evk4_bright_track.gif): | |
| 32 consecutive windows from the magnitude-7.3 EVK4 recording. Tracker IoU exceeded raw | |
| detector IoU in 27 of the 32 displayed labelled windows. | |
| - [DAVIS SAOCOM-1B tracker-smoothing animation](../competition_ready_improved_cpu_student/visualizations/animations/davis_saocom_tracker_smoothing.gif): | |
| 29 consecutive windows at native 346×260 resolution. Tracker IoU was higher in 15 of the | |
| displayed labelled windows. | |
| - [DVX Stars3 tracker-smoothing animation](../competition_ready_improved_cpu_student/visualizations/animations/dvx_stars_tracker_smoothing.gif): | |
| 30 consecutive cluttered windows. Tracker IoU was higher in 18 displayed windows. | |
| - [DVX Thuraya3 low-SNR animation](../competition_ready_improved_cpu_student/visualizations/animations/dvx_thuraya_low_snr.gif): | |
| 33 consecutive windows in the hardest dim-target regime; 30 contain GT, and tracker IoU was | |
| higher in 19 of those labelled frames. | |
|  | |
|  | |
|  | |
|  | |
| These clips were deliberately selected to explain behavior. Their per-frame counts are qualitative | |
| descriptions of the shown intervals, not new aggregate accuracy estimates or blind evidence. The | |
| animation manifest records the exact sequence/window ranges, promoted cache, final policy, and GIF | |
| hashes needed to reproduce them. | |
| ## 16. Docker and reproducibility architecture | |
| The final Linux/amd64 image is designed to be inspected and run without network access. It | |
| contains: | |
| - the exact student checkpoint and SHA-256 sidecar; | |
| - the compact model definition and preprocessing code; | |
| - sensor parameters, tracker policy, and writer implementation; | |
| - pinned CPU runtime dependencies; | |
| - the official evaluator; | |
| - an automatic `CMD ["sh", "run.sh"]` entrypoint; | |
| - a model-structure document checked against the baked runtime settings. | |
| At runtime it: | |
| 1. reads `*.npy` sequences from `/OrbitSight_dataset`; | |
| 2. refuses ambiguous sensors or duplicate sequence names; | |
| 3. verifies the checkpoint hash; | |
| 4. prewarms Numba and one model window; | |
| 5. resets all state at each sequence boundary; | |
| 6. writes `<sequencename>.txt` files with exactly nine tab-separated columns; | |
| 7. runs the bundled evaluator; | |
| 8. writes `Evaluation_Metrics.xlsx`; | |
| 9. publishes only the known files under `/work/SIMER/DDMMYYYY`. | |
| The final improved archive SHA-256 is: | |
| `3b2db2bdf3cc7b1d331f8fa6dc8dcc23281bb831177ff07c75f273eb0b819eb5` | |
| Archive validation checks the Linux/amd64 manifest, layer diff IDs, automatic command, | |
| mount contract, output schema, model hash, runtime dependencies, calibration constants, | |
| tracker settings, and baked model-structure text. | |
| ## 17. Evidence boundaries and scientific honesty | |
| Three different claims must not be confused: | |
| 1. **Reproducibility:** the same files and code reproduce a measured local result. | |
| 2. **Score maximization:** a setting scores better on recordings already examined. | |
| 3. **Blind generalization:** a frozen setting performs better on recordings that had no role | |
| in training, tuning, diagnostics, or selection. | |
| OrbitSight has strong evidence for the first claim. The routed hybrid and final calibration | |
| also have evidence for the second. There are no unused local recordings, so neither provides | |
| a new blind confirmation. The organizer's unseen finalist dataset is the correct test of the | |
| third claim. | |
| The leakage-free aggregate removes the known DAVIS session overlap, but the remaining files | |
| have still been used for diagnostics. “Leakage-free” here means free of that known training- | |
| session contamination, not untouched by all development decisions. | |
| ## 18. Why the final student is the recommended submission | |
| The heavy hybrid has the best local accuracy but fails the CPU requirement by an order of | |
| magnitude. V5 has a cleaner scientific history but its corrected full-tail latency exceeded | |
| 40 ms on some held-out regions, and its local mAP is below the compact student's. The compact | |
| student offers: | |
| - one small checkpoint rather than a routed pair; | |
| - 135,111 parameters; | |
| - full 640×640 small-target localization; | |
| - recurrent temporal integration; | |
| - sensor-aware but computationally trivial calibration; | |
| - a local worst-sequence p95 near 21 ms; | |
| - F1 near the V5 scientific baseline; | |
| - better local mAP than V5; | |
| - a smaller, clearer, CPU-only Docker delivery. | |
| This is a practical systems decision rather than a claim that the student is universally | |
| more accurate than every predecessor. | |
| ## 19. Recommended next work | |
| If new recordings become available, the next work should be data-centered rather than another | |
| threshold sweep: | |
| 1. Freeze the current and rollback Docker hashes before receiving new data. | |
| 2. Keep new labels hidden while both versions produce predictions. | |
| 3. Compare both on exactly the same capture-disjoint recordings. | |
| 4. Require improved F1, no mAP regression, no sensor-family collapse, and p95 below 40 ms. | |
| 5. Treat that dataset as spent after results are opened. | |
| For model research, the highest-value directions are: | |
| - collect more low-SNR DVX positives and representative empty star fields; | |
| - train a better proposal-ranking objective for hard negatives; | |
| - explore calibrated simulation only after matching real event statistics; | |
| - profile OpenVINO/BF16/INT8 only with strict output-parity and accuracy gates; | |
| - measure container startup, file reading, formatting, workbook creation, and disk output on | |
| the organizer's exact i9-12900H hardware. | |
| ## 20. Conclusion | |
| OrbitSight progressed from an honest sub-0.1 baseline to a complete real-time CPU system by | |
| treating data contracts, localization, temporal reasoning, deployment, and evaluation as one | |
| problem. V5 established a credible scientific detector. V7 contributed a superior EVK4 | |
| specialist. Their hybrid demonstrated the best available local accuracy but exposed an | |
| unavoidable latency gap. Selective teacher–student distillation transferred much of that | |
| behavior into a model small enough for real-time CPU inference. Final calibration improved | |
| the frozen student's operating point without retraining. | |
| The current submission is therefore the product of the whole development path—not merely the | |
| last threshold and tracker adjustment. Its strongest qualities are balanced performance, | |
| transparent evidence, explicit limitations, and a reproducible offline delivery contract. | |
| ## Selected references | |
| 1. G. Gallego et al., “Event-based Vision: A Survey,” IEEE TPAMI, 2020. | |
| DOI: 10.1109/TPAMI.2020.3008413. | |
| 2. S. Afshar et al., “Event-based Object Detection and Tracking for Space Situational | |
| Awareness,” arXiv:1911.08730, 2019. | |
| 3. N. O. Ralph et al., “Astrometric Calibration and Source Characterisation of the Latest | |
| Generation Neuromorphic Event-based Cameras for Space Imaging,” arXiv:2211.09939, 2022. | |
| 4. Y. Alkendi et al., “Dynamic Space Object Detection with Neuromorphic Vision Sensors,” | |
| European Conference on Space Debris, 2025. | |
| 5. E. Perot et al., “Learning to Detect Objects with a 1 Megapixel Event Camera,” NeurIPS, 2020. | |
| 6. M. Gehrig and D. Scaramuzza, “Recurrent Vision Transformers for Object Detection with | |
| Event Cameras,” CVPR, 2023. | |
| 7. X. Shi et al., “Convolutional LSTM Network: A Machine Learning Approach for Precipitation | |
| Nowcasting,” NeurIPS, 2015. | |
| 8. Z. Tian et al., “FCOS: Fully Convolutional One-Stage Object Detection,” ICCV, 2019. | |
| 9. Z. Ge et al., “YOLOX: Exceeding YOLO Series in 2021,” arXiv:2107.08430, 2021. | |
| 10. T.-Y. Lin et al., “Focal Loss for Dense Object Detection,” ICCV, 2017. | |
| 11. Z. Zheng et al., “Distance-IoU Loss: Faster and Better Learning for Bounding Box | |
| Regression,” AAAI, 2020. DOI: 10.1609/aaai.v34i07.6999. | |
| 12. G. Hinton, O. Vinyals, and J. Dean, “Distilling the Knowledge in a Neural Network,” | |
| arXiv:1503.02531, 2015. | |
| 13. Y. Zhang et al., “ByteTrack: Multi-Object Tracking by Associating Every Detection Box,” | |
| ECCV, 2022. | |
| ## Primary local evidence | |
| - `docs/technical_report.md` | |
| - `docs/v7_audit_status.md` | |
| - `notes.md` | |
| - `artifacts/v7_gates/official_test_confirmation/epoch33-paired-official-once-v3/report.json` | |
| - `artifacts/cpu_student/hybrid_distilled_student_eval_latency_summary.json` | |
| - `artifacts/cpu_student/hybrid_distilled_student_full_cpu_latency.json` | |
| - `artifacts/cpu_student/post_training_latency_neutral/promotion.json` | |
| - `artifacts/cpu_student/post_training_latency_neutral/tracker_live_accuracy.json` | |
| - `artifacts/cpu_student/post_training_latency_neutral/tracker_live_latency.json` | |