نوع مقاله : مقاله پژوهشی
عنوان مقاله English
نویسندگان English
Traffic crashes are a major public health and urban safety challenge in Tehran, where vulnerable road users account for a substantial share of annual fatalities. Using fatal crash records from 2017 to 2025, this study applied descriptive statistics, the chi-square test, and Cramér's V to assess associations among categorical variables; eight supervised machine learning algorithms to predict deceased road-user type and road type; and K-Means clustering to identify latent patterns. Annual crash counts increased markedly. Pedestrians and motorcyclists constituted a large share of vulnerable road user fatalities, and highways were the most common crash location. Based on consistency across cross-validation folds, XGBoost performed best in predicting deceased road-user type (accuracy = 0.8091). For road type, severe class imbalance was addressed by comparing 14 imbalance-handling methods according to recall for the minority class, secondary streets (7.3% of records). Only an adjusted random forest substantially improved secondary-street recall (from 0 to 0.3333) without a marked loss in overall accuracy (0.4704). Clustering identified four profiles: Cluster 0, characterized by a notable pedestrian presence; Cluster 1, by a higher share of motorcyclists and highway crashes; Cluster 2, by pedestrian predominance and the highest proportion of hit-and-run incidents; and Cluster 3, the largest, by motorcyclist predominance. These differences underscore the need for interventions tailored to each group to improve traffic safety management in Tehran.
کلیدواژهها English