Abstract— This study presents a semi-automated and reproducible pipeline for extracting high-quality, morphology-preserved single RBC images from peripheral blood smears of Thai patients. The dataset includes six clinically confirmed hematological conditions: Iron Deficiency Anemia, Thalassemia Trait, Hb H Disease, Thalassemia Hb E Disease, Symptomatic Thalassemia Hb E Disease, and Homozygous Hb E. Whole-slide images were acquired using an Aperio AT2 scanner at 40× magnification and processed through a Python-based workflow integrating mean shift filtering, adaptive thresholding, contour-based segmentation, and edge-contact rejection to exclude incomplete or touching cells. Isolated RBCs were filtered by size and overlaid onto standardized black backgrounds ranging from 32×32 to 1024×1024 pixels. The proposed pipeline generated a curated dataset of 12,229 single-cell RBC images, with the 128×128 pixel group identified as the most stable and representative resolution across all disease categories. Baseline CNN experiments demonstrate that the extracted images are learnable and structurally consistent, confirming dataset usability without claiming optimized or clinically deployable performance. By providing a region-specific RBC dataset derived from a Thai cohort with a high prevalence of hemoglobinopathies, together with a transparent extraction framework and baseline CNN validation, this work establishes a solid foundation for future CNN-based RBC morphology analysis and population-specific digital hematopathology research.
Keywords: Red blood cell segmentation; Blood smear analysis; Image preprocessing; Medical image dataset
DOI: https://doi.org/10.5455/jjee.204-1751358445

![Scopus®_151_PNG-300x86[1]](https://jjee.ttu.edu.jo/wp-content/uploads/2024/03/Scopus®_151_PNG-300x861-1.png)
