Proactive forecasting of student academic outcomes via Random Forest and cumulative learning data
Abstract
This study develops a Random Forest-based predictive framework to forecast course outcomes for Information Technology students at Binh Duong University using institutional academic records. Key features such as course codes, instructor identifiers, and encoded student IDs were combined with demographic and performance data through a standardised preprocessing pipeline that included feature engineering, normalisation, and class-imbalance handling. The final model achieved an accuracy of 0.98, a ROC AUC of 0.982, and a recall of 0.84 for the “Fail” class, demonstrating strong predictive power and practical value for early risk detection. Feature-importance analysis confirmed the dominant influence of course- and instructor-related variables alongside cumulative student histories, ensuring interpretability. By providing reliable and transparent predictions, the framework supports academic advisers in identifying at-risk students and designing timely interventions, highlighting the potential of institutional data for enhancing educational decision-making.
Keywords:
Random Forest, Data Mining, student evaluation, Educational data miningDOI:
https://doi.org/10.31276/VJSTE.2025.0052Classification number
1.2
Downloads
Published
Received 2 July 2025; revised 8 August 2025; accepted 5 October 2025




