NobleBlocks

CMKL University

UniversityBangkok, Bangkok, Thailand

Research output, citation impact, and the most-cited recent papers from CMKL University (Thailand). Aggregated across the NobleBlocks index of 300M+ scholarly works.

Total works
50
Citations
98
h-index
5
i10-index
1
Also known as
CMKL Universityมหาวิทยาลัยซีเอ็มเคแอล

Top-cited papers from CMKL University

An alternative approach to ontology-based curriculum development in higher education
Pattamaporn Piriyapongpipat, Sally E Goldin, Nadh Ditcharoen
2024· Smart Learning Environments8doi:10.1186/s40561-024-00307-8

Abstract Global trends in higher education emphasize the development of curricula that offer greater responsiveness to learners. Creating flexible and responsive curricula will require additional support systems for curriculum management. The first step toward sustainably developing this kind of system is to represent essential curricular information in a way that allows sharing common components across various work processes within the educational environment. The current research implements a new approach for representing curriculum components, by systematically analyzing external source data to extract the basic knowledge, skills and dependencies which then become objects into an ontology. The resulting ontology should act as a computationally-accessible model of the curriculum with sufficient information and usable quality. This paper describes a trial implementation of our approach using actual curriculum documents. Results from performance metrics and expert evaluation validate the proposed strategy and suggest that the approach is feasible for real-world practice.

Dependable Sensing System for Pig Farming
Sripong Ariyadech, Amelie Bonde, Orathai Sangpetch, Woranun Woramontri +4 more
20195doi:10.1109/gciot47977.2019.9058398

We have deployed our smart sensing system in a real commercial farm complex for at least 17 months. One of critical factors to success of the sensing system deployment is the resilience and fault tolerance of the system in a harsh environment with unreliable infrastructure and limited access. Power interruptions and intermittent connectivity is not uncommon. Sensors must work even submerging in animal excretion. Correct and continuous streams of sensor data is essential to our smart farming analytics. To make our sensing system sustain such challenges, we have designed and implemented the system with a capability of self-rejuvenation to ensure system liveness. We also equip it with our anomaly detection system to examine small sensor connectivity logs in order to identify potential faulty or deteriorating sensors or external event abnormality with minimal manual intervention.

Thai-Dialect: Low Resource Thai Dialectal Speech to Text Corpora
Artit Suwanbandit, Jaturong Chitiyaphol, Sutthinan Chuenchom, Kanyarat Kwiecien +4 more
20234doi:10.1109/asru57964.2023.10389792

We release a speech-to-text benchmark dataset containing 10 Thai dialects that cover different regions of Thailand. Our corpora consists of the standard dialect, Thai-central (THA); the northern dialects (Khummuang (NOD), Nan (KHB) and Yno (YNO)); the northeastern dialects (Korat (TTS), Khmer (KXM) and Laos (TTS)); and the southern dialects (Krabi (SOU), Pattani (MFA) and Phangnga (SOU)). All transcriptions are based on the Thai writing standard. We constructed baseline models by fine-tuning from self-supervised pre-trained models. Results show that multilingual/multidialectal systems outperform monolingual ones, and different dialect combinations can affect the performance of multilingual/multidialectal training.

GaussianSlicer: Efficient Surface Reconstruction from Cross-sectional Slices with Gaussian Splatting
Yuhu Guo, Chenghao Qian, Yuhong Mo, Akkarit Sangpetch
20253doi:10.1109/icassp49660.2025.10890834

In this work, we present GaussianSlicer, an efficient plane-based Gaussian Splatting framework for reconstructing geometry from cross-sectional slices input. Unlike previous methods that rely on computationally intensive geometry-based or grid-based implicit techniques, which struggle with complex cases (e.g., sparse slices, multi-hole geometries). GaussianSlicer enables parallel optimization without requiring any prior constraints. Our method begins by initializing planar Gaussians on each slice and optimizing their layout to obtain accurate geometry representation. To align the Gaussian splats, we introduce a geometric regularization that promotes surface smoothness and ensures consistency in global topology. Our system enables accurate 3D reconstruction from sparse, irregular, multi-label slices with high computational efficiency. Experimental results show that, on average, our method is 45.62% faster and achieves a 21.91% improvement in Chamfer Distance (CD) outperforming state-of-the-art methods on the collected dataset.

Optimal Antenna Slot Design for Hepatocellular Carcinoma Microwave Ablation Using Multi-Objective Fuzzy Decision Making
Petch Nantivatana, Pattarapong Phasukkit, Supan Tungjitkusolmun, Keerati Chayakulkheeree
2020· International journal of intelligent engineering and systems3doi:10.22266/ijies2020.1031.05

This paper presents the optimal antenna slot position and sizing (OASPS), for microwave (MW) hepatocellular carcinoma (HCC) treatment, using multi-objective fuzzy decision making (MOFDM).In the optimal design problem formulation, the multi-objective for, (1) achievement for a near-spherical zone for ablation defined by axial ratio (AR), (2) the minimum reflection coefficient (S 11 ), and (3) maximum volume of destroying (VD), are desired.The finite element method (FEM) is used for MW distribution simulation for coordination analysis with the proposed MOFDM based OASPS.The investigation on the design results shown that the temperature distribution of the proposed MOFDM for OASPS is in the more controllable shape comparing to the existing design, due to the sharp beam of temperature distribution to the target at the front side of slot with the less temperature distribution at the rare side of slot.The results also show that the proposed MOFDM for OASPS for OASPS design is easier to manage with the less effect to the rare side of the slot.Moreover, the proposed MOFDM for OASPS resulted in the temperature distribution shape closer to round shape than the existing design.The MOFDM for OASPS is tested on designing of the antenna build up from the coaxial cable.The coaxial cable used is Semi-rigid 141 (RG402 M17/130-RG402 Copper Jacket) including Inner conductor, Dielectric, and Outer conductor.In the simulation, the slot distance from conductor end (L ts ), representing the slot position, is varied as 2.3mm, 3.3mm, 4.3mm, 5.3mm, 6.3mm, 7.3mm, and 8.3mm.Meanwhile, the slot size (W d ), is varied as 1mm, 2mm, 3mm, 4mm, 5mm, and 6mm.The input power (P i ) used is 50 W with the duration of 300 seconds.The comparison of obtained by different L ts and W d , is investigated.The simulation results show that the proposed MOFDM based OASPS can efficiently and effectively provide the near-spherical zone of ablation result with simultaneously trade-off between S 11 minimization and VD maximization.

Comparison of 4-dimensional variational and ensemble optimal interpolation data assimilation systems using a Regional Ocean Modeling System (v3.4) configuration of the eddy-dominated East Australian Current system
Colette Kerry, Moninya Roughan, Shane Richard Keating, David E. Gwyther +3 more
2024· Geoscientific model development2doi:10.5194/gmd-17-2359-2024

Ocean models must be regularly updated through the assimilation of observations (data assimilation) in order to correctly represent the timing and locations of eddies. Since initial conditions play an important role in the quality of short-term ocean forecasts, an effective data assimilation scheme to produce accurate state estimates is key to improving prediction. Western boundary current regions, such as the East Australia Current system, are highly variable regions, making them particularly challenging to model and predict. This study assesses the performance of two ocean data assimilation systems in the East Australian Current system over a 2-year period. We compare the time-dependent 4-dimensional variational (4D-Var) data assimilation system with the more computationally efficient, time-independent ensemble optimal interpolation (EnOI) system, across a common modelling and observational framework. Both systems assimilate the same observations: satellite-derived sea surface height, sea surface temperature, vertical profiles of temperature and salinity (from Argo floats), and temperature profiles from expendable bathythermographs. We analyse both systems' performance against independent data that are withheld, allowing a thorough analysis of system performance. The 4D-Var system is 25 times more expensive but outperforms the EnOI system against both assimilated and independent observations at the surface and subsurface. For forecast horizons of 5 d, root-mean-squared forecast errors are 20 %–60 % higher for the EnOI system compared to the 4D-Var system. The 4D-Var system, which assimilates observations over 5 d windows, provides a smoother transition from the end of the forecast to the subsequent analysis field. The EnOI system displays elevated low-frequency (>1 d) surface-intensified variability in temperature and elevated kinetic energy at length scales less than 100 km at the beginning of the forecast windows. The 4D-Var system displays elevated energy in the near-inertial range throughout the water column, with the wavenumber kinetic energy spectra remaining unchanged upon assimilation. Overall, this comparison shows quantitatively that the 4D-Var system results in improved predictability as the analysis provides a smoother and more dynamically balanced fit between the observations and the model's time-evolving flow. This advocates the use of advanced, time-dependent data assimilation methods, particularly for highly variable oceanic regions, and motivates future work into further improving data assimilation schemes.

HoloGrad: A Holographic Health Information Platform for Patient Education - Delivering Personalized Genetic Information and Counseling to Users
S. Chan-Bormei, Hossein Miri
20242doi:10.1145/3657547.3657560

Personalized and user-centered health information systems are indispensable in healthcare: (1) they provide customized medical information and advice that can positively influence individual treatment or therapy success; (2) they can improve doctor-patient communication towards enhanced patient comprehension. A major challenge during doctor-patient discussions is the communication and visualization of information, procedures, events, or outcomes. This is typically achieved through conversations, videos, slides, or simply drawing on paper. However, in the area of medical genetics, 3D visualizations could greatly assist in informing and educating patients. They could also convey personalized information more easily and more effectively than mere 2D imagery, helping patients better understand their situations and complex genetic information, leading to improved support in pre- and post-counseling sessions. Such impactful outcomes also have pedagogical implications for educational practices, as the need for individualizing knowledge and information transmission to patients is aptly applicable to education too. In this paper, we report on the design, development, and implementation of an immersive, personalized, user-centered, and interactive health information platform that employs a holographic visualization interface to allow 3D interactions with mixed reality views within the field of medical genetics, in order to support the presentation and delivery of specialized information to patients in a clear and comprehensible manner. Our proposed approach has the potential to enhance information comprehension, increase information retention, aid in doctor-patient communication, and shorten the length of genetic counseling sessions, by presenting 3D visualizations on mixed really headsets, such as HoloLens and MetaQuestPro. Patients will be able to see in 3D what a geneticist or genetic counselor is trying to demonstrate, with the help of holograms and virtual simulations. Therefore, viewing and manipulating complicated genetic structures as well as explaining the procedures involved in genetic screening and genetic testing will become clearer and more understandable. The platform, dubbed HoloGrad, is essentially a novel, interactive, and customizable holographic genetic visualization tool where medical professionals as well as medical translators and nurses can perceive and examine genetic data and procedures, and visually explore 3D virtual perspectives using intuitive hand gestures, voice commands, hand-controllers, and gesture-controlled keyboard holograms.

Comparison of 4-Dimensional Variational and Ensemble Optimal Interpolation data assimilation systems using a Regional Ocean Modelling System (v3.4) configuration of the eddy-dominated East Australian Current System
Colette Kerry, Moninya Roughan, Shane Richard Keating, David E. Gwyther +3 more
20232doi:10.5194/egusphere-2023-2355

Abstract. Ocean models must be regularly updated through the assimilation of observations (data assimilation) in order to correctly represent the timing and locations of eddies. Since initial conditions play an important role in the quality of short-term ocean forecasts, an effective data assimilation scheme to produce accurate state estimates is key to improving prediction. Western boundary current regions, such as the East Australia Current system, are highly variable regions making them particularly challenging to model and predict. This study assesses the performance of two ocean data assimilation systems in the East Australian Current system over a two year period. We compare the time-dependent 4-Dimensional Variational (4D-Var) data assimilation system with the more computationally-efficient, time-independent Ensemble Optimal Interpolation (EnOI) system, across a common modelling and observational framework. Both systems assimilate the same observations including: satellite-derived sea-surface height, sea-surface temperature, vertical profiles of temperature and salinity (from Argo floats), and temperature profiles from eXpendable Bathy-Thermographs. We analyse both systems' performance against independent data that is withheld allowing a thorough analysis of system performance. The 4D-Var system is 25 times more expensive but outperforms the EnOI system against both assimilated and independent observations at the surface and subsurface. For forecast horizons of 5-days Root-mean-squared forecast errors are 20–60 % higher for the EnOI system compared to the 4D-Var system. The 4D-Var system, which assimilates observations over 5-day windows, provides a smoother transition from the end of the forecast to the subsequent analysis field. The EnOI system displays elevated low frequency (>1 day), surface intensified variability in temperature, and elevated kinetic energy at length scales less than 100 km at the beginning of the forecast windows. The 4D-Var system displays elevated energy in the near-inertial range throughout the water column, with the wavenumber kinetic energy spectra remaining unchanged upon assimilation. Overall, this comparison shows quantitatively that the 4D-Var system results in improved predictability as the analysis provides a smoother and more dynamically-balanced fit between the observations and the model's time-evolving flow. This advocates the use of advanced, time-dependent data assimilation methods, particularly for highly variable oceanic regions, and motivates future work into further improving data assimilation schemes.

GPU Performance Tuning and Power Efficiency on the DGX A100 Cluster
Khanin Udomchoksakul, Orathai Sangpetch, Akkarit Sangpetch
20222doi:10.1109/cloudcom55334.2022.00033

The complexity of current Deep learning has been growing rapidly nowadays. Such advancement allows various organizations such as private sectors and government to leverage intelligent systems on their use cases. High Performance Computing (HPC) infrastructure nowadays has pivoted to GPU-oriented systems, enabling developers and researchers to train complex models with large datasets unlike conventional clusters equipped only with CPU cores. However, focus on power efficiency on the HPC system has not been prevalent especially on the new system such as DGX A100 that does not have datapoints on how GPUs consumed power. Even though such HPC cluster can be powerful, always allowing it to run at the maximum capacity results to financial cost to the HPC provider at the end. Therefore, for any organization providing the system, it is crucial for them to balance the cluster capabilities while maintaining overall power consumption which can potentially be costly in the long term. This paper reveals A100 GPU metrics that are relevant to Power usage and explains GPU profiling applied to Deep learning workload on the cluster, saving up to 32% of the power usage while compromising only 11.5% of training time compared to a default profile. Then, the paper investigates literature review that could be learned further adopted to the current system at CMKL university as the next milestone.

Improved Joint Estimation for Body-Mounted Motion Capture Sensors Using Human Kinematics Prior Knowledge
Shaun Stevens, Paulo Alonso Gaona-García, Hyong Kim
2022· 2022 IEEE Sensors2doi:10.1109/sensors52175.2022.9967322

Measurement uncertainty is affected by several factors, including sensor resolution. In circumstances where uncertainty varies across features (e.g., distance to transducer), confidence in sensor results is particularly affected, reducing their applicability in applications such as joint estimation through body sensors. In this manuscript, a model for joint estimation based on millimeter wave point cloud data and prior knowledge of human body structure is discussed. The proposed model uses structure and kinematic information about the measured entity to reduce measurement uncertainty, even when the expected uncertainty bounds change according to distance to the transducer. This is achieved by augmenting a sensor's processing stage with additional estimation constraints using prior knowledge, and using intersection of error bounds to improve estimated values. For a body-mounted motion tracking sensor performing joint estimation, our model is compared to empirical results of two 3D convolutional neural networks, one associated with a head-mounted millimeter wave sensor, and the other associated with a body-mounted millimeter wave sensor. The joint estimation model is found to outperform the neural-network based estimator in both mounting scenarios, resulting in reduced estimation error as confirmed by an external sensor that provides ground truth.

PIWIMS: Physics Informed Warehouse Inventory Monitory via Synthetic Data Generation
João Diogo Falcão, Prabh Simran Singh Baweja, Yi Wang, Akkarit Sangpetch +3 more
20212doi:10.1145/3460418.3480415

State-of-the-art camera-based deep learning methods for inventory monitoring tend to fail to generalize across different domains due to the high variance of scene settings. Large amounts of human labor are required to label and parameterize the models, making a real-world deployment impractical. In a third-party warehouse setting, supervised learning approaches are either too costly and/or inaccurate to deploy due to the need for human labor to address the diverse set of environmental factors (i.e, lighting conditions, product motion, deployment limitations).

PEX: Privacy-Preserved, Multi-Tier Exchange Framework for Cross Platform Virtual Assets Trading
Akkarit Sangpetch, Orathai Sangpetch
20202doi:10.1109/ccnc46108.2020.9045515

In traditional virtual asset trading market, several risks, e.g. scams, cheating users, and market reach, have been pushed to users (sellers/buyers). Users need to decide who to trust; otherwise, no business. This fact impedes the growth of virtual asset trading market. In the past few years, several virtual asset marketplaces have embraced blockchain and smart contract technology to alleviate such risks, while trying to address privacy and scalability issues. To attain both speed and non-repudiation property for all transactions, existing blockchain-based exchange systems still cannot fully accomplish. In real-life trading, users use traditional contract to provide non-repudiation to achieve accountability in all committed transactions, so-called thorough non-repudiation. This is essential when dispute happens. To achieve similar thorough non-repudiation as well as privacy and scalability, we propose PEX, Privacy-preserved, multi-tier EXchange framework for cross platform virtual assets trading. PEX creates a smart contract for each virtual asset trading request. The key to address the challenges is to devise two-level distributed ledgers with two different types of quorums where one is for public knowledge in a global ledger and the other is for confidential information in a private ledger. A private quorum is formed to process individual smart contract and record the transactions in a private distributed ledger in order to maintain privacy. Smart contract execution checkpoints will be continuously written in a global ledger to strengthen thorough non-repudiation. PEX smart contract can be executed in parallel to promote scalability. PEX is also equipped with our reputation-based network to track contribution and discourage malicious behavior nodes or users, building healthy virtual asset ecosystem.

Building RSSI-based Indoor Positioning Fingerprint Maps using Android-based Coordination
Lapat Nakpaen, Prab Wongsekleo, Panarat Cherntanomwong, Charnon Pattiyanon
20241doi:10.1109/isai-nlp64410.2024.10799385

Indoor positioning systems (IPS) have emerged as a critical technology for location-based applications. Developing IPS system is challenging since technologies for outdoor positioning seem to be limited in indoor environment. Fingerprinting is a technique to build an offline map and compare the current location with it. While fingerprinting remains a popular technique for indoor positioning, its reliance on extensive manual data collection is a significant challenge. These data points can be the Received Signal Strength Indicator (RSSI) of the Wi-Fi signal or signals from the triangulation of Bluetooth/cellular beacons. However, the conventional grid-based fingerprint technique is facing challenges when the target area is being large. This research proposes an automated approach to gathering Wi-Fi RSSI data for building indoor positioning maps using the Android-based triangulated coordination. Our method demonstrates a substantial reduction in data collection time (79%) compared to traditional grid-based techniques. The resulting dataset effectively supports machine learning models for indoor positioning, achieving a Mean Distance Error (MDE) of less than 2 meters different.

Optimizing YOLOv8 for Efficient Tomato Recognition in Greenhouse Environments Using Drone Imagery
Oleg Cohan Shovkovyy, Hossein Miri
20241doi:10.1109/asiancomnet63184.2024.10811019

This study explores the application and fine-tuning of You Only Look Once (YOLOv8) models for real-time tomato recognition using drone imagery in greenhouse environments, with a focus on practical optimization strategies. Our evaluation of YOLO’s speed, robustness, and adaptability revealed that varying batch sizes and epochs had minimal impact on performance. Notably, the YOLOv8n model matched the performance of the YOLOv8x model while reducing training time by up to 60 times. Further fine-tuning identified the final learning rate (lrf) and dataset annotation quality as critical factors for model performance. Optimizing the lrf and enhancing dataset annotations significantly improved accuracy, underscoring their importance in effective YOLO model deployment. Our results demonstrate YOLOv8’s superiority over YOLOv5, with the optimized YOLOv8n model being ready for deployment in future tomato recognition tasks, paving the way for more efficient agricultural monitoring. This work provides valuable insights into object detection and offers practical guidance for researchers addressing similar challenges.

PHYOT: Physics-Informed Object Tracking in Surveillance Cameras
Kawisorn Kamtue, José M. F. Moura, Orathai Sangpetch, Paulo Alonso Gaona-García
20241doi:10.1109/icassp48485.2024.10448150

While deep learning has been very successful in computer vision, real world operating conditions such as lighting variation, background clutter, or occlusion hinder its accuracy across several tasks. Prior work has shown that hybrid models—combining neural networks and heuristics/algorithms—can outperform vanilla deep learning for several computer vision tasks, such as classification or tracking.We consider the case of object tracking, and evaluate a hybrid model (PhyOT) that conceptualizes deep neural networks as "sensors" in a Kalman filter setup, where prior knowledge, in the form of Newtonian laws of motion, is used to fuse sensor observations and to perform improved estimations. Our experiments combine three neural networks, performing position, indirect velocity and acceleration estimation, respectively, and evaluate such a formulation on two benchmark datasets: a warehouse security camera dataset that we collected and annotated and a traffic camera open dataset.Results suggest that our PhyOT can track objects in extreme conditions that the state-of-the-art deep neural networks fail while its performance in general cases does not degrade significantly from that of existing deep learning approaches. Results also suggest that our PhyOT components are generalizable and transferable.

Internet of Wearables: Fog Extrapolation for Reduced Data Collection and Expanded Capture Volume in Real-Time Motion Capture Edge Devices
Shaun Stevens, Paulo Alonso Gaona-García, Hyong Kim
20221doi:10.1109/cloudcom55334.2022.00030

The range of applications that make up the Internet-of-Things ecosystem continues to grow. New opportunities present themselves along with new design challenges concerning the efficiency, portability, and processing capabilities of future Internet-of-Things systems. Thus, improving the operational metrics of individual Internet-of-Things devices, particularly across the edge and fog layers, is of paramount importance.In this manuscript, we present an approach for decreasing data collection at the edge, thus reducing form factor and power consumption of edge devices. This is particularly relevant for our application of interest, wearable motion capture, where human comfort and operational longevity are of prime importance. Our approach extrapolates from reduced edge data by leveraging prior physiological knowledge of the captured entity at the computational stage in the fog. By delegating computation to the fog, we also demonstrate the possibility for expanded capture volumes (operational areas) for future wearable motion capture systems and motion capture systems in general.Our approach, when prototyped on a millimeter wave sensor edge device and two fog node platforms of different processing tiers, shows that prior knowledge can facilitate a reduction in capture data dimensionality (and an associated decrease in power consumption) with little to no accuracy degradation, when compared to a more data-intensive edge system (Microsoft Kinect).

Regional and spatial dependence of poverty factors in Thailand, and its use in Bayesian hierarchical regression analysis
Irving Gómez-Méndez, Chainarong Amornbunchornvej
2026· Statistical Journal of the IAOSdoi:10.1177/18747655261416694

Poverty in Thailand shows strong spatial dependence that existing administrative boundaries fail to capture, leading to policies that overlook local socioeconomic realities. This study proposes a data-driven regionalization framework to infer geographically coherent “policy regions” that better represent poverty dynamics. Using household-level data from the Thai People Map and Analytics Platform (TPMAP), we analyze spatial autocorrelation across multiple poverty factors through Moran’s statistics and principal component analysis, followed by spatially constrained hierarchical clustering to delineate coherent regions. Bayesian hierarchical and geographically weighted regression models are then employed to examine how education influences household income at provincial, regional, and national levels. Our results identified six regions that reflect more accurately poverty structures than official divisions. Northern and Northeastern Thailand emerge as the regions most affected by low education, income, and savings, while Central Thailand shows higher inequality. The inferred regions demonstrate that spatially contiguous provinces often share similar socioeconomic structures, suggesting that policy targeting should align with these patterns rather than provincial borders. Our findings provide a quantitative foundation for evidence-based regional planning, enabling policymakers to design differentiated yet regionally coordinated interventions. The approach illustrates how spatial statistical modeling can bridge the gap between data analysis and effective poverty-alleviation policy.

Quantum–Classical Hybrid LSTM Model for PM2.5 Forecasting in Bangkok
Rajasurang Wongkrasaemongkol
2025doi:10.1109/incit66780.2025.11276072

Air pollution is a serious environmental and health concern in Bangkok. This research proposes a hybrid quantum-classical model for forecasting fine particulate matter ($\text{P M}_{\text{2} \cdot \text{5}}$) by using Quantum Long Short-Term Memory (QLSTM) architecture that embeds a variational quantum circuit within the recurrent gating pathway to construct a compact, nonlinear feature map in Hilbert space. Daily concentrations of$\text{PM}_{2 \cdot 5}, \text{PM}_{10}, \mathrm{O}_{3}, \text{NO}_{2}, \text{SO}_{2}$, and CO are used to train and evaluate the models on a held-out test set. The QLSTM improves from the LSTM baseline in both RMSE and MSE evaluation metrics. The method utilizes quantum superposition and entanglement to enhance feature representation, enabling faster convergence and reduced initial error. The findings underscore the integration of quantum computing into machine learning models to achieve high predictive accuracy with compact parameterizations for environmental time-series forecasting.

Multi-Dimensional Road Extraction from Medium Resolution Imagery using U-Net and YOLO Concepts
Chutchatut Sutichavengkul, Sally E Goldin
2025doi:10.1109/icitee66631.2025.11338258

Digital geographic databases support vital applications such as navigation, disaster response and urban planning. Accurate road network information is especially important for many of these uses. Recognizing that remotely sensed data offer broad geographic coverage and temporal currency, recent research has successfully applied deep learning (DL) techniques to extract roads from satellite imagery. However, almost all these studies have used high resolution images which are expensive to acquire and to process. Furthermore, these DL models produce segmentation maps that indicate only the presence or absence of a road in each output pixel. We introduce a novel methodology which uses more economical medium resolution satellite imagery and produces a rich, multi-dimensional result. Our approach integrates a U-Net segmentation backbone with enhancements inspired by single-stage object detection models to produce a multi-channel output comprising a pixel-level probability map, sub-pixel offset coordinates for refined road alignment and orientation angles that capture local directional trends. The framework generates independent vector centerline proposals without merging them into existing maps. While it does not currently perform integration, these proposals can be used as input to downstream workflows, such as manual validation or OSM updates, for more reliable large-scale mapping.

Replacing YOLOv8n with YOLOv12n in Greenhouse Automation: A Post-Thesis Evaluation Using the Unified Detection Score (UDS)
Oleg Cohan Shovkovyy, Hossein Miri
2025doi:10.1109/iscit67082.2025.11231599

This study presents a post-thesis evaluation of object detection models for greenhouse automation, extending previous efforts that applied YOLOv8n to classify tomato ripeness stages in Thai agricultural settings. The original system integrated drone-based image acquisition with IoT-driven irrigation control, demonstrating that compact AI models can operate effectively even in resource-limited agricultural environments. However, its evaluation relied primarily on conventional metrics, such as mAP@50 and precision, which may overlook nuanced model behavior in real-world conditions. To address this gap, we introduce the Unified Detection Score (UDS) - a composite metric designed to provide a holistic and interpretable measure of model quality. Using the re-annotated tomato ripeness dataset from the original research (“Precision Agriculture in Thailand: Merging Drone-Assisted Object Recognition with IoT,” 2024), this study compares five YOLOv12 variants (n, s, m, l, x) against YOLOv8n and YOLOv5. Among them, YOLOv12n achieved the highest UDS of 93.3, outperforming YOLOv8n (91.9) and YOLOv5 (82.7), while maintaining a compact model size of 5.4 MB and a fast-training time of 0.52 hours. Larger YOLOv12 variants offered no significant performance advantage; in fact, YOLOv12x, despite its scale, underperformed with a UDS of 90.9 and a size of 119.5 MB. Further experiments confirmed that the original YOLOv8n training configuration (50 epochs, batch size 8) remains optimal for YOLOv12n, with extended training routines yielding negligible gains. Optimizer screening validated AdamW as the most effective choice, and hyperparameter tuning revealed that a decay of 0.0009 and a momentum of 0.975 offer the best balance of precision and stability. These findings highlight YOLOv12n as a practical and efficient successor to YOLOv8n for greenhouse automation tasks. By introducing the UDS framework and offering detailed comparisons across multiple architectures, this study contributes both a novel evaluation methodology and practical insights to help advance scalable, precision agriculture solutions.