AraDigitalContent77: A Novel Arabic Intent Classification Dataset for Digital Content Platforms via Cross-Domain Schema Adaptation
Keywords:
Arabic NLP, Intent Classification, Schema Adaptation, Low-Resource Datasets, AraBERT, Digital Content Platforms, Synthetic Data GenerationAbstract
Intent classification underpins conversational and recommendation systems, yet Arabic remains underserved by domain-specific benchmarks outside banking and finance. This paper introduces AraDigitalContent77, an Arabic intent classification dataset for digital content platforms, including streaming, podcasts, music, news, and audiobooks. The Banking77 schema was adapted through an Action-Object-Goal framework, producing 77 intents composed of 7 direct transfers, 34 analogical transfers, and 36 newly defined intents; 36 banking-specific intents were removed. Utterances were created through a five-stage hybrid pipeline combining human-written seeds, controlled large language model generation, automated quality filtering, selective human review, template-pattern transformation, and intentional linguistic-noise injection. The resulting corpus contains 13,090 utterances evenly distributed across 77 intents, with 170 utterances per intent, no exact duplicate sentences, and no missing values. A fine-tuned AraBERT classifier was evaluated using stratified 70/15/15 training, validation, and independent test splits. The final test set produced 96.59% accuracy and a macro F1-score of 96.61%. These results establish a strong within-dataset baseline while avoiding direct superiority claims across benchmarks that use different domains and evaluation protocols. The dataset and reproducible splits are publicly released under a CC BY 4.0 license.