Fine-Tuning YOLOv8 for Top-View Vehicle Detection: Dataset Preparation, Training, and Performance Analysis for Traffic Surveillance

Authors

  • Muhammad Zain Ul Abdin Department of Computer Science, NCBA&E, Multan, 60000, Pakistan

DOI:

https://doi.org/10.66108/mna.v5i01.88

Keywords:

YOLOv8, Traffic Surveillance, Vehicle Detection, Top-View Vehicle Detection, Object Detection

Abstract

Traffic surveillance systems require a very high level of accuracy and real-time vehicle detection so it can be used to enable intelligent transportation and city management applications. The study proposes a detailed study on how to fine-tune the YOLOv8 object detector to top-view aerial vehicle detector in traffic monitoring sites. The model was trained using a domain-specific dataset of 626 annotated aerial images of vehicles, which with the use of transfer learning delay the optimization of detection precision and recall. The training was done with a batch size of 32, the resolution of the input normalized to 640×640 pixels, and a 100-epoch run in its full Tesla P100 GA. The model is good at finding different types of vehicles in different traffic situations since it has a high mean Average Precision (mAP50) of 97.5% and a strong mAP50-95 of 74.2%, as well as precision and recall rates over 91%. The ability to draw conclusions within seconds (about 4 milliseconds per picture) renders the technique even more applicable to the live traffic surveillance systems. Integration with Weights & Biases made it possible to keep an eye on all aspects of the training process, which made sure that convergence was stable and the model could be used in other situations. The results show how important it is to use specialized datasets and fine-tune them to improve systems for detecting aerial vehicles for smart traffic management. The main argument of the present study is that a domain-specific highly accurate aerial vehicle detector can be obtained using a rather small dataset with the use of efficient fine-tuning of YOLOv8. Such research provides a computationally low and efficient model compared to the literature available, which is constituted by huge data sets, or a complex architecture that could be utilized in real time and in traffic monitoring systems.

Downloads

Download data is not yet available.

References

Ren, S., He, K., Girshick, R., & Sun, J. (2016). Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE transactions on pattern analysis and machine intelligence, 39(6), 1137-1149. 10.1109/TPAMI.2016.2577031

Ramos, L. T., & Sappa, A. D. (2025). A decade of you only look once (yolo) for object detection: A review. IEEE Access, 13, 192747-192794. 10.1109/ACCESS.2025.3630988

Cheng, G., Han, J., & Lu, X. (2017). Remote sensing image scene classification: Benchmark and state of the art. Proceedings of the IEEE, 105(10), 1865-1883. 10.1109/JPROC.2017.2675998

Girshick, R., Donahue, J., Darrell, T., & Malik, J. (2014). Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 580-587). 10.1109/CVPR.2014.81

Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C. Y., & Berg, A. C. (2016, September). Ssd: Single shot multibox detector. In European conference on computer vision (pp. 21-37). Cham: Springer International Publishing. https://doi.org/10.1007/978-3-319-46448-0_2

Redmon, J., & Farhadi, A. (2017). YOLO9000: better, faster, stronger. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 7263-7271). https://doi.org/10.48550/arXiv.1612.08242

Redmon, J., Divvala, S., Girshick, R., & Farhadi, A. (2016). You only look once: Unified, real-time object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 779-788). https://doi.org/10.48550/arXiv.1506.02640

Wu, Y., Guan, X., Zhao, B., Ni, L., & Huang, M. (2023). Vehicle detection based on adaptive multimodal feature fusion and cross-modal vehicle index using RGB-T images. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 16, 8166-8177. 10.1109/JSTARS.2023.3294624

Dikbayir, H. S., & BÜLBÜL, H. Ï. (2020, December). Deep learning based vehicle detection from aerial images. In 2020 19th IEEE International Conference on Machine Learning and Applications (ICMLA) (pp. 956-960). IEEE. 10.1109/ICMLA51294.2020.00155

Dai, J., Li, Y., He, K., & Sun, J. (2016). R-fcn: Object detection via region-based fully convolutional networks. Advances in neural information processing systems, 29. https://doi.org/10.48550/arXiv.1605.06409

Pan, S. J., & Yang, Q. (2009). A survey on transfer learning. IEEE Transactions on knowledge and data engineering, 22(10), 1345-1359. 10.1109/TKDE.2009.191

Widayani, A., Putra, A. M., Maghriebi, A. R., Adi, D. Z. C., & Ridho, M. H. F. (2024). Review of application YOLOv8 in medical imaging. Physics Letters, 5(1). 10.20473/iapl.v5i1.57001

Everingham, M., Eslami, S. A., Van Gool, L., Williams, C. K., Winn, J., & Zisserman, A. (2015). The pascal visual object classes challenge: A retrospective. International journal of computer vision, 111(1), 98-136. https://doi.org/10.1007/s11263-014-0733-5

Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., & Zitnick, C. L. (2014). Microsoft COCO: Common Objects in Context. Computer Vision – ECCV 2014, 740–755. https://doi.org/10.1007/978-3-319-10602-1_48

Additional Files

Published

2026-03-16

How to Cite

Muhammad Zain Ul Abdin. (2026). Fine-Tuning YOLOv8 for Top-View Vehicle Detection: Dataset Preparation, Training, and Performance Analysis for Traffic Surveillance. Machines and Algorithms, 5(01), 3–12. https://doi.org/10.66108/mna.v5i01.88

Issue

Section

Articles

Categories