Detection of Fraud in Financial Statements Based on Managerial Tone Analysis, Financial Reporting Complexity, and Anomaly Detection Algorithms
The present study aimed to identify fraudulent financial reporting based on managerial tone analysis, financial reporting complexity, and anomaly detection algorithms among companies listed on the Tehran Stock Exchange. This study was conducted using a quantitative applied research design with a correlational and explanatory approach based on machine learning and textual analysis techniques. The statistical population consisted of companies listed on the Tehran Stock Exchange between 2018 and 2024, from which 186 firms were selected using purposive screening criteria, resulting in 1302 firm-year observations. Data were collected from audited annual reports, financial statements, explanatory notes, and management discussion sections. Managerial tone was measured through sentiment analysis using natural language processing methods, while financial reporting complexity was assessed through readability indicators, report length, and disclosure structure measures. Fraudulent reporting risk was identified using abnormal accrual indicators and anomaly detection techniques including Isolation Forest, Local Outlier Factor, One-Class Support Vector Machine, Random Forest, and Autoencoder Neural Networks. Data analysis was conducted using Python, SPSS-27, and RapidMiner software through descriptive statistics, correlation analysis, logistic regression, ROC curve analysis, and machine learning model evaluation indicators including accuracy, precision, recall, F1-score, and AUC. The findings demonstrated that positive managerial tone (β = 0.31, p < 0.001), negative managerial tone (β = 0.36, p < 0.001), financial reporting complexity (β = 0.39, p < 0.001), Fog readability index (β = 0.28, p < 0.001), abnormal accruals (β = 0.47, p < 0.001), and Isolation Forest anomaly scores (β = 0.51, p < 0.001) significantly predicted fraudulent financial reporting. Correlation analysis revealed significant positive relationships among fraud risk, reporting complexity, abnormal accruals, and anomaly detection indicators (p < 0.01). Among the evaluated algorithms, the Autoencoder Neural Network demonstrated the highest predictive performance with an accuracy of 0.93 and AUC of 0.96, followed by Isolation Forest with an accuracy of 0.91 and AUC of 0.94. Machine learning-based anomaly detection techniques significantly outperformed traditional logistic regression models in identifying suspicious financial reporting patterns. The results indicate that integrating managerial tone analysis, financial reporting complexity indicators, and anomaly detection algorithms substantially improves the identification of fraudulent financial statements. The findings highlight the importance of combining textual analysis, artificial intelligence, and forensic accounting techniques within modern auditing and financial supervision systems. Furthermore, the superior performance of deep learning and anomaly detection models suggests that advanced computational technologies can significantly enhance fraud detection accuracy and support auditors and regulators in identifying hidden manipulation patterns within corporate financial disclosures.