Methodology behind EKa'a
Our approach to extracting and analyzing contemporary Arabic song lyrics to identify societal portrayals.
Overview
Eka'a applies Anmat's text-analysis workflow to cultural material rather than news coverage. The project explores how contemporary Arabic song lyrics encode recurring social portrayals and thematic associations across a requested corpus.
Data sources
The dataset is assembled by scraping a requested artist's complete discography from a dedicated Arabic lyrics site. Because the corpus is produced on demand, its boundaries depend on the research question, the selected artist or artists, and what remains accessible at collection time.
Method
We crawl the selected lyric pages, normalize the collected text, and segment each song into two-line couplets as the unit of analysis. We then apply Association Rule Mining (ARM) using the Apriori algorithm to detect and measure which terms, motifs, and portrayals appear together repeatedly across these couplets.
Limitations & caveats
The corpus depends on what is published online and what can be safely collected from the lyric source. It is not a complete map of Arabic music, and recurring textual patterns still require human interpretation before making cultural claims.
Access model
This project is generated on request rather than released as a fixed public dataset. That keeps the extraction tightly scoped to a clear research question and makes the collection choices easier to document for collaborators.
