Search

Results for "AraNizer"

ArabianGPT: Native Arabic GPT-based Large Language Model

arXiv · Feb 23

The paper introduces ArabianGPT, a suite of transformer-based language models designed specifically for Arabic, including versions with 0.1B and 0.3B parameters. A key component is the AraNizer tokenizer, tailored for Arabic script's morphology. Fine-tuning ArabianGPT-0.1B achieved 95% accuracy in sentiment analysis, up from 56% in the base model, and improved F1 scores in summarization. Why it matters: The models address the gap in native Arabic LLMs, offering better performance on Arabic NLP tasks through tailored architecture and tokenization.

Fanar aims to advance Arabic presence in digital space| Gulf Times - Gulf Times

QCRI · Jun 14

The provided article content is missing, preventing a factual summary of its details. Information regarding Fanar's initiative to advance Arabic presence in the digital space could not be extracted. Specific actions, partnerships, or funding related to this endeavor are not available. Why it matters: Without the article content, the significance of Fanar's potential contributions to Arabic digital presence cannot be evaluated.

AraNet: A Deep Learning Toolkit for Arabic Social Media

arXiv · Dec 30

Researchers introduce AraNet, a deep learning toolkit for Arabic social media processing. The toolkit uses BERT models trained on social media datasets to predict age, dialect, gender, emotion, irony, and sentiment. AraNet achieves state-of-the-art or competitive performance on these tasks without feature engineering. Why it matters: The public release of AraNet accelerates Arabic NLP research by providing a comprehensive, deep learning-based tool for various social media analysis tasks.

AraGPT2: Pre-Trained Transformer for Arabic Language Generation

arXiv · Dec 31

The paper introduces AraGPT2, a suite of pre-trained transformer models for Arabic language generation, with the largest model (AraGPT2-mega) containing 1.46 billion parameters. Trained on a large Arabic corpus of internet text and news, AraGPT2-mega demonstrates strong performance in synthetic news generation and zero-shot question answering. To address the risk of misuse, the authors also released a discriminator model with 98% accuracy in detecting AI-generated text. Why it matters: This release of both the model and discriminator fills a critical gap in Arabic NLP and encourages further research and applications in the field.