MelenL commited on
Commit
30ddf64
·
verified ·
1 Parent(s): 1eaffa3

Add Custom Inception-style CNN project card

Browse files
Files changed (1) hide show
  1. README.md +76 -0
README.md CHANGED
@@ -1,3 +1,79 @@
1
  ---
2
  license: mit
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  license: mit
3
+ language:
4
+ - en
5
+ pipeline_tag: image-classification
6
+ library_name: keras
7
+ metrics:
8
+ - accuracy
9
+ tags:
10
+ - computer-vision
11
+ - satellite-imagery
12
+ - remote-sensing
13
+ - xview
14
+ - cnn
15
+ - inception
16
+ - tensorflow
17
+ - trained-from-scratch
18
  ---
19
+
20
+ # Custom Inception-style CNN for Satellite Image Classification
21
+
22
+ A custom convolutional neural network trained from scratch to classify satellite image crops into 13 xView categories. Developed as a Deep Learning course project at Universidad Politécnica de Madrid (UPM), this architecture uses parallel convolutional branches to extract features at multiple spatial scales.
23
+
24
+ ## Architecture
25
+
26
+ - **Input:** 128 × 128 RGB image crops.
27
+ - **Output:** softmax probabilities over 13 classes.
28
+ - **Parameters:** 4,529,325, as recorded in the notebook.
29
+ - **Feature extractor:** convolutional stem followed by seven custom Inception-style modules.
30
+ - **Parallel branches:** 1 × 1 convolution; 1 × 1 followed by 3 × 3 convolution; 1 × 1 followed by two 3 × 3 convolutions; and max pooling followed by 1 × 1 convolution.
31
+ - **Classification head:** global average pooling, dropout (0.4), and a dense softmax layer.
32
+ - **Framework:** TensorFlow / Keras.
33
+
34
+ This is a custom Inception-inspired architecture, not the standard InceptionV3 model. It classifies individual image crops rather than detecting objects in full satellite scenes.
35
+
36
+ ## Training
37
+
38
+ The notebook uses Adam with an initial learning rate of 0.001 and categorical cross-entropy with label smoothing of 0.1. Training is configured for up to 50 epochs with a batch size of 64. The saved training log identifies epoch 47 as the best epoch by validation accuracy.
39
+
40
+ ## Results
41
+
42
+ The notebook compares three custom CNN architectures on the same validation split:
43
+
44
+ | Architecture | Validation accuracy |
45
+ | --- | --- |
46
+ | ResNet-style | 18.67% |
47
+ | VGG-style | 68.53% |
48
+ | **Inception-style** | **72.69%** |
49
+
50
+ The selected Inception-style model also achieved **75.63% macro recall** and **75.48% macro precision** in the recorded validation evaluation.
51
+
52
+ These results come from the original experiments in `CNN Best model.ipynb`. They refer to the course's 13-class classification setup, not the full xView object detection benchmark.
53
+
54
+ ## Classes
55
+
56
+ Cargo plane, small car, bus, truck, motorboat, fishing vessel, dump truck, excavator, building, helipad, storage tank, shipping container, and pylon.
57
+
58
+ ## Project materials
59
+
60
+ - Training and evaluation notebook: `CNN Best model.ipynb`.
61
+ - Project report: `Report_ImageRecognitionAndObjectDetectiononthexViewSatelliteDataset.pdf`.
62
+
63
+ The notebook documents the architecture definitions, training experiments, confusion matrices, and per-class evaluation. Results from any subsequent training run should be evaluated independently.
64
+
65
+ ## Authors
66
+
67
+ Melen Laclais, Léo Lamy, and Adrián García-Pozuelo Fornieles.
68
+
69
+ ## License and attribution
70
+
71
+ The MIT license designation applies to original project code only. Third-party code and course materials retain their respective terms.
72
+
73
+ xView imagery and annotations remain under **CC BY-NC-SA 4.0**, including any dataset images reproduced in notebooks or the report.
74
+
75
+ - [xView dataset](https://xviewdataset.org/)
76
+ - [Official xView dataset license terms](https://challenge.xviewdataset.org/rules)
77
+ - [CC BY-NC-SA 4.0](https://creativecommons.org/licenses/by-nc-sa/4.0/)
78
+
79
+ Dataset reference: Lam et al., *xView: Objects in Context in Overhead Imagery* (2018).