60 datasets found
  1. Bengali Sentiment Dataset

    • kaggle.com
    zip
    Updated Jul 25, 2020
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Nuhash Afnan (2020). Bengali Sentiment Dataset [Dataset]. https://www.kaggle.com/nuhashafnan/pseudolabel
    Explore at:
    zip(631287 bytes)Available download formats
    Dataset updated
    Jul 25, 2020
    Authors
    Nuhash Afnan
    Description

    Dataset

    This dataset was created by Nuhash Afnan

    Contents

  2. h

    BLUGE-bengali-sentiment-classification

    • huggingface.co
    Updated Jul 16, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Nahid Hossain (2026). BLUGE-bengali-sentiment-classification [Dataset]. https://huggingface.co/datasets/nahid-hub/BLUGE-bengali-sentiment-classification
    Explore at:
    Dataset updated
    Jul 16, 2026
    Authors
    Nahid Hossain
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    BLUGE-TSC: Bangla Sentiment Classification

    BLUGE-TSC is a meticulously curated and cleaned Bangla Ternary Sentiment Classification dataset, one of the 7 tasks in BLUGE (Bengali Language UnderstandinG Evaluation), a balanced benchmark for evaluating Bengali natural language understanding. See the full BLUGE collection for all 7 tasks, and the B-CORE pretraining corpus and BnLM model suite released alongside it.

      Dataset Description
    

    This task classifies Bangla text… See the full description on the dataset page: https://huggingface.co/datasets/nahid-hub/BLUGE-bengali-sentiment-classification.

  3. BanglaMUSE: Bangla Text–Audio Sentiment Dataset

    • kaggle.com
    zip
    Updated Jan 8, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Yasout141516 (2026). BanglaMUSE: Bangla Text–Audio Sentiment Dataset [Dataset]. https://www.kaggle.com/datasets/yasout141516/banglamuse-bangla-textaudio-sentiment-dataset/code
    Explore at:
    zip(306676882 bytes)Available download formats
    Dataset updated
    Jan 8, 2026
    Authors
    Yasout141516
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    BanglaMUSE is a multimodal Bangla sentiment dataset containing aligned text–audio pairs designed for research in sentiment analysis, speech processing, and multimodal learning for low-resource languages.

    The dataset includes 1,000 Bangla sentences, evenly balanced between positive (500) and negative (500) sentiment classes. The sentences represent natural, everyday Bangla language usage and were manually curated and validated to ensure clear sentiment polarity.

    Each sentence is recorded by four native Bangla speakers (two female and two male), resulting in 4,000 speech recordings in total. All speakers recorded the same set of sentences, enabling controlled analysis of speaker variability while preserving identical textual content. Audio samples are provided in MP3 format, recorded in controlled indoor environments, and manually verified for quality and alignment.

    The dataset is distributed with a unified metadata.csv file that links sentence identifiers, sentiment labels, speaker information, and relative audio paths. BanglaMUSE supports tasks such as multimodal sentiment classification, sentiment-aware speech recognition, audio–text alignment, and speaker-independent modeling.

  4. m

    Data Set For Sentiment Analysis On Bengali News Comments

    • data.mendeley.com
    Updated Sep 15, 2019
    + more versions
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Md Akhter Uz Zaman Ashik Chowdhury (2019). Data Set For Sentiment Analysis On Bengali News Comments [Dataset]. http://doi.org/10.17632/n53xt69gnf.2
    Explore at:
    Dataset updated
    Sep 15, 2019
    Authors
    Md Akhter Uz Zaman Ashik Chowdhury
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    This is a data set of Sentiment Analysis On Bangla News Comments where every data was annotated by three different individuals to get three different perspectives and based on the majorities decisions the final tag was chosen. This data set contains 13802 data in total.

  5. m

    Bengali Political Sentiment Analysis Dataset

    • data.mendeley.com
    Updated Oct 2, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Adib Mahmud (2025). Bengali Political Sentiment Analysis Dataset [Dataset]. http://doi.org/10.17632/x5yc4m5yg2.2
    Explore at:
    Dataset updated
    Oct 2, 2025
    Authors
    Adib Mahmud
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    This dataset comprises 3,290 Bengali political comments sourced from social media platforms, news comment sections, and online political discussions, specifically curated for sentiment analysis research in Bengali NLP. The corpus provides a comprehensive resource for training and evaluating sentiment classification models within the political domain. The dataset features 3,290 instances distributed across five sentiment classes with excellent balance (variance <8%): Very Negative (675, 20.5%), Negative (663, 20.2%), Neutral (626, 19.0%), Very Positive (664, 20.2%), and Positive (662, 20.1%). Stored in Excel format with two columns containing Bengali political comments (Unicode text) and corresponding sentiment labels, the dataset maintains high quality with no missing values and verified annotations. Comment lengths average 83 characters, ranging from 11 to 398 characters. The collection encompasses diverse political discourse including government policies and governance, electoral processes and democracy, political parties and leadership dynamics, social and economic issues, current affairs and political events, along with public opinion and citizen responses to political developments. This dataset serves multiple research purposes, including Bengali sentiment analysis model development and benchmarking, political discourse analysis and opinion mining, natural language processing research for low-resource languages, cross-lingual sentiment analysis studies, social media analytics for Bengali content, multi-class text classification research, and comparative political sentiment studies across different linguistic and cultural contexts.

  6. SentiFive: A Multi-Class Bengali Dataset

    • kaggle.com
    zip
    Updated Nov 11, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Amer Mahbub (2025). SentiFive: A Multi-Class Bengali Dataset [Dataset]. https://www.kaggle.com/datasets/amermahbub01/sentifive
    Explore at:
    zip(1566621 bytes)Available download formats
    Dataset updated
    Nov 11, 2025
    Authors
    Amer Mahbub
    License

    Attribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
    License information was derived automatically

    Description

    SentiFive is a multi-class Bengali sentiment analysis dataset consisting of 31,411 YouTube comments, manually annotated into five sentiment categories: Strongly Negative, Weakly Negative, Neutral, Weakly Positive, and Strongly Positive. The dataset is designed to support research in fine-grained sentiment classification and low-resource language processing.

    Unlike previous Bengali sentiment datasets that focus on binary or ternary sentiment, SentiFive enables more nuanced modeling of user opinions expressed in informal social media contexts. Data were collected from a diverse set of YouTube videos, covering topics such as news, entertainment, and politics.

    M. A. Rahman, A. Mahbub, B. N. Paul, P. Bhattacharjee and M. A. -U. -Z. Ashik, "SentiFive: A Multi-Class Bengali Dataset for Sentiment Analysis," 2025 IEEE 7th International Conference on Sustainable Technologies For Industry 5.0 (STI), Dhaka, Bangladesh, 2025, pp. 1-6, doi: 10.1109/STI69347.2025.11367591. keywords: {Deep learning;Sentiment analysis;Video on demand;Social networking (online);Bidirectional long short term memory;Web sites;Reliability;Fifth Industrial Revolution;Standards;Software development management;SentiFive;5-Class Sentiment;Baseline Evaluation;LSTM;BiLSTM;Bangla Natural Language Processing (BNLP);Sentiment Analysis (SA)},

  7. Bengali Sentiment Analysis Dataset

    • kaggle.com
    zip
    Updated Sep 27, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Rhs Liza (2025). Bengali Sentiment Analysis Dataset [Dataset]. https://www.kaggle.com/datasets/rhsliza/bengali-sentiment-analysis-dataset
    Explore at:
    zip(16210 bytes)Available download formats
    Dataset updated
    Sep 27, 2025
    Authors
    Rhs Liza
    License

    https://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/

    Description

    Dataset

    This dataset was created by Rhs Liza

    Released under CC0: Public Domain

    Contents

  8. S

    BanglaMUSE-VID: A Bangla Video-Based Sentiment Analysis Dataset

    • scidb.cn
    Updated Jul 1, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Md. Masudul Islam (2026). BanglaMUSE-VID: A Bangla Video-Based Sentiment Analysis Dataset [Dataset]. http://doi.org/10.57760/sciencedb.35376
    Explore at:
    CroissantCroissant is a format for machine-learning datasets. Learn more about this at mlcommons.org/croissant.
    Dataset updated
    Jul 1, 2026
    Dataset provided by
    Science Data Bank
    Authors
    Md. Masudul Islam
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    The BanglaMUSE-VID dataset is publicly available through the Science Data Bank (SciDB) and provides a synchronized multimodal benchmark for Bangla sentiment analysis. The repository contains 1,000 manually curated Bangla sentences with balanced sentiment labels (500 positive and 500 negative) and 5,000 corresponding face-visible MP4 video recordings captured by five native Bangla speakers using different smartphone devices. The dataset comprises both textual and video modalities, enabling research in multimodal sentiment analysis, visual speech recognition, cross-modal representation learning, and video-grounded language understanding for low-resource Bangla NLP applications. The repository is openly accessible and designed for future extensibility through the inclusion of additional speakers, sentiment categories, and domain-specific content.

  9. m

    Bangla Online Comments Dataset

    • data.mendeley.com
    Updated Jan 28, 2021
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Md Faisal Ahmed (2021). Bangla Online Comments Dataset [Dataset]. http://doi.org/10.17632/9xjx8twk8p.1
    Explore at:
    Dataset updated
    Jan 28, 2021
    Authors
    Md Faisal Ahmed
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    The total amount of collected comments is 44001. The dataset aims to differentiate whether a comment is a bully expression or not with the help of Natural Language Processing and to what extent it is improper if it is an inappropriate comment. The comments are labeled with different categories of harassment with the help of experts and consensus.

  10. m

    Data from: ANUBHUTI: A COMPREHENSIVE CORPUS FOR SENTIMENT ANALYSIS IN BANGLA...

    • data.mendeley.com
    Updated Jan 19, 2026
    + more versions
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Swastika Kundu (2026). ANUBHUTI: A COMPREHENSIVE CORPUS FOR SENTIMENT ANALYSIS IN BANGLA REGIONAL LANGUAGES [Dataset]. http://doi.org/10.17632/mjxwby94yw.3
    Explore at:
    Dataset updated
    Jan 19, 2026
    Authors
    Swastika Kundu
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    ANUBHUTI, a comprehensive dataset consisting of 2,500 sentences manually translated from standard Bangla into four major regional dialects—Mymensingh, Noakhali, Sylhet, and Chittagong. The dataset predominantly features political and religious content, reflecting the contemporary socio-political landscape of Bangladesh, alongside neutral texts to maintain balance. Each sentence is annotated using a dual annotation scheme: (i) multiclass thematic labeling categorizes sentences as Political, Religious, or Neutral, and (ii) multilabel emotion annotation assigns one or more emotions from Anger, Contempt, Disgust, Enjoyment, Fear, Sadness, and Surprise.

  11. Bangla Sentiment Analysis + Microblog Posts

    • kaggle.com
    zip
    Updated May 15, 2022
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Wasifa Chowdhury (2022). Bangla Sentiment Analysis + Microblog Posts [Dataset]. https://www.kaggle.com/wchowdhu/bengali-sentiment-analysis-microblog-posts
    Explore at:
    zip(74749 bytes)Available download formats
    Dataset updated
    May 15, 2022
    Authors
    Wasifa Chowdhury
    Description

    Contains training and test sets as well as the manually created Twitter-specific lexicons for the following paper,

    Performing Sentiment Analysis in Bangla Microblog Posts

    The resources are made available to encourage more research in Bangla sentiment classification and other NLP tasks.

    Please cite the paper if you use the dataset or lexicon

  12. Code Mixed Sentiment [Bangla-English-Hindi]

    • kaggle.com
    zip
    Updated Oct 2, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Md Nishat Raihan (2023). Code Mixed Sentiment [Bangla-English-Hindi] [Dataset]. https://www.kaggle.com/datasets/mdnishatraihan/code-mixed-sentiment-bangla-english-hindi
    Explore at:
    zip(31311608 bytes)Available download formats
    Dataset updated
    Oct 2, 2023
    Authors
    Md Nishat Raihan
    License

    Attribution-NoDerivs 4.0 (CC BY-ND 4.0)https://creativecommons.org/licenses/by-nd/4.0/
    License information was derived automatically

    Description

    This is a dataset for the sentiment analysis task. It contains 100k code mixed data. The languages are Bangla-English-Hindi.

    Dataset Generation:

    Initially, we select the Amazon Review Dataset as our base data, referenced from Ni et al. (2019)**[1]**. We randomly extract 100,000 instances from this dataset. The original labels in this dataset are ratings, scaled from 1 to 5. For our specific task, we categorize them into Positive (rating > 3), Neutral (rating = 3), and Negative (rating < 3), ensuring a balanced number of instances for each label. To generate the synthetic Code-mixed dataset, we apply two distinct methodologies: the Random Code-mixing Algorithm by Krishnan et al. (2021)**[2]** and r-CM by Santy et al. (2021)**[3]**.

    Class Distribution:

    For train.csv:

    LabelCountPercentage
    Negative2000033.33%
    Neutral2000033.33%
    Positive1999933.33%

    For dev.csv:

    LabelCountPercentage
    Neutral666733.34%
    Positive666733.34%
    Negative666633.33%

    For test.csv:

    LabelCountPercentage
    Negative666733.34%
    Positive666733.34%
    Neutral666633.33%

    Cite our Paper:

    If you utilize this dataset, kindly cite our paper.

    @article{raihan2023mixed, title={Mixed-Distil-BERT: Code-mixed Language Modeling for Bangla, English, and Hindi}, author={Raihan, Md Nishat and Goswami, Dhiman and Mahmud, Antara}, journal={arXiv preprint arXiv:2309.10272}, year={2023} }

    References

    [1]: Ni, J., Li, J., & McAuley, J. (2019). Justifying recommendations using distantly-labeled reviews and fine-grained aspects. In Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP) (pp. 188-197).

    [2]: Krishnan, J., Anastasopoulos, A., Purohit, H., & Rangwala, H. (2021). Multilingual code-switching for zero-shot cross-lingual intent prediction and slot filling. arXiv preprint arXiv:2103.07792.

    [3]: Santy, S., Srinivasan, A., & Choudhury, M. (2021). BERTologiCoMix: How does code-mixing interact with multilingual BERT? In Proceedings of the Second Workshop on Domain Adaptation for NLP (pp. 111-121).

  13. m

    Data from: ANUBHUTI: A COMPREHENSIVE CORPUS FOR SENTIMENT ANALYSIS IN BANGLA...

    • data.mendeley.com
    Updated Jun 26, 2025
    + more versions
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Swastika Kundu (2025). ANUBHUTI: A COMPREHENSIVE CORPUS FOR SENTIMENT ANALYSIS IN BANGLA REGIONAL LANGUAGES [Dataset]. http://doi.org/10.17632/mjxwby94yw.1
    Explore at:
    Dataset updated
    Jun 26, 2025
    Authors
    Swastika Kundu
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    ANUBHUTI, a comprehensive dataset consisting of 2,000 sentences manually translated from standard Bangla into four major regional dialects—Mymensingh, Noakhali, Sylhet, and Chittagong. The dataset predominantly features political and religious content, reflecting the contemporary socio-political landscape of Bangladesh, alongside neutral texts to maintain balance. Each sentence is annotated using a dual annotation scheme: (i) multiclass thematic labeling categorizes sentences as Political, Religious, or Neutral, and (ii) multilabel emotion annotation assigns one or more emotions from Anger, Contempt, Disgust, Enjoyment, Fear, Sadness, and Surprise.

  14. Bangla Sentiment DataSet

    • kaggle.com
    zip
    Updated Aug 15, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    M Arman (2025). Bangla Sentiment DataSet [Dataset]. https://www.kaggle.com/mithilaarman2/bangla-sentiment-dataset
    Explore at:
    zip(821003 bytes)Available download formats
    Dataset updated
    Aug 15, 2025
    Authors
    M Arman
    Description

    Dataset

    This dataset was created by M Arman

    Contents

  15. Motamot Bengali Political Sentiment Analysis

    • kaggle.com
    zip
    Updated Aug 18, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Mukaffi Moin (2025). Motamot Bengali Political Sentiment Analysis [Dataset]. https://www.kaggle.com/datasets/mukaffimoin/motamot-bengali-political-sentiment-analysis/code
    Explore at:
    zip(5515870 bytes)Available download formats
    Dataset updated
    Aug 18, 2025
    Authors
    Mukaffi Moin
    License

    MIT Licensehttps://opensource.org/licenses/MIT
    License information was derived automatically

    Description

    Motamot: Bengali Political Sentiment Analysis Dataset

    📖 Overview

    Motamot is a Bengali political sentiment analysis dataset containing 7,058 labeled data points. Each entry is annotated with Positive or Negative sentiment, specifically tailored for analyzing political discourse in the Bengali language.

    This dataset supports Natural Language Processing (NLP) research, with applications in sentiment classification, political opinion mining, and benchmarking pre-trained and large language models (LLMs) for low-resource languages.

    📊 Dataset Statistics

    SplitTotalPositiveNegative
    Train564733062341
    Test706413293
    Validation705413292
    Total705841322926

    📜 Citation

    If you use this dataset, please cite the following paper:

    @INPROCEEDINGS{10752197,
     author={Johora Faria, Fatema Tuj and Moin, Mukaffi Bin and Mumu, Rabeya Islam and Alam Abir, Md Mahabubul and Alfy, Abrar Nawar and Alam, Mohammad Shafiul},
     booktitle={2024 IEEE Region 10 Symposium (TENSYMP)}, 
     title={Motamot: A Dataset for Revealing the Supremacy of Large Language Models Over Transformer Models in Bengali Political Sentiment Analysis}, 
     year={2024},
     pages={1-8},
     keywords={Sentiment analysis;Analytical models;Accuracy;Voting;Large language models;Transformers;Market research;Few shot learning;Portals;IEEE Regions;Political Sentiment Analysis;Pre-trained Language Models;Large Language Models;Gemini 1.5 Pro;GPT 3.5 Turbo;Zero-shot Learning;Fewshot Learning;Low-resource Language},
     doi={10.1109/TENSYMP61132.2024.10752197}
    }
    

    📂 Dataset Structure

    Motamot/
    │
    ├── train.csv    # Training set (5,647 instances)
    ├── test.csv     # Test set (706 instances)
    ├── validation.csv  # Validation set (705 instances)
    

    Each file contains:

    • text → Bengali political statement
    • label → Sentiment category (Positive or Negative)

    🧪 Baseline Results

    🔹 Comparative Analysis of Pre-trained Language Models

    ModelAccuracyPrecisionRecallF1-Score
    BanglaBERT0.82040.82220.82040.8203
    Bangla BERT Base0.68030.69070.68120.6833
    DistilBERT0.63200.63580.63200.6317
    mBERT0.64270.64960.64280.6153
    sahajBERT0.67080.67910.67090.6707

    🔹 Comparative Analysis of Large Language Models (Few-shot & Zero-shot)

    Model (LLM)MetricZero-shot5-shot10-shot15-shot
    GPT 3.5 TurboAccuracy0.85000.89000.91330.9400
    Precision0.84670.88670.92000.9467
    Recall0.85330.89260.90790.9342
    F1-Score0.84950.88960.91390.9404
    Gemini 1.5 ProAccuracy0.86080.89810.92000.9633
    Precision0.89310.88460.93330.9667
    Recall0.84770.92050.90910.9603
    F1-Score0.86980.90220.92110.9635

    📬 Contact Information

    For questions, collaborations, or inquiries:

  16. m

    Advancing Bengali NLP for Sentiment and Emotion Dataset

    • data.mendeley.com
    Updated Jan 27, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Rownuk Ara Rumy (2025). Advancing Bengali NLP for Sentiment and Emotion Dataset [Dataset]. http://doi.org/10.17632/kztpv8g89p.1
    Explore at:
    Dataset updated
    Jan 27, 2025
    Authors
    Rownuk Ara Rumy
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    The dataset consists of 34,812 Bengali posts and comments sourced from Facebook, Twitter, and Instagram, Bengali news portals and literature. Techniques employed in data acquisition included data scraping from social media accounts through API and scraping only text data from websites. Microblogs consist of posts and comments from platforms like Facebook, Twitter, and Instagram, which allow for the capture of informal and emotionally rich text. Newspaper and magazine articles provide formal, sentiment-related information through opinions. Online literature, including Bengali novels, poems, and blogs, incorporates semantic relationships and linguistic nuances. Text data is collected from public sources through automated scripts. We used selenium scripts, created using the Python programming language. We used APIs to obtain structured social media data. Additionally, we complied with the requirements of privacy, data collection, and ethics.It contains 5 Emotion and 5 Sentiment class. For emotion "Creepy" being the most frequent emotion with 12,000 entries, followed by "Unbiased" with 8,500 entries, "Joyful" with 7,500 entries, "Bullying" with 4,000 entries, and "Surprise" with 2,500 entries. On the other hand, for sentiment "Negative" being the most frequent with 8,000 entries, followed by "Neutral" with 7,000 entries, "Strongly Negative" with 6,800 entries, "Positive" with 5,500 entries, and "Strongly Positive" with 4,500 entries in that order.

  17. M

    BanglaBlend: A Large-Scale Nobel Dataset of Bangla Sentences Categorized by...

    • datasetcatalog.nlm.nih.gov
    Updated Dec 9, 2024
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    ayman, Umme Ayman; Saha, Chayti; Mawa, Zannatul (2024). BanglaBlend: A Large-Scale Nobel Dataset of Bangla Sentences Categorized by Saint(Sadhu) and Common(Cholito) Form of Bengali Language [Dataset]. https://datasetcatalog.nlm.nih.gov/dataset?q=0001420662
    Explore at:
    Dataset updated
    Dec 9, 2024
    Authors
    ayman, Umme Ayman; Saha, Chayti; Mawa, Zannatul
    Description

    This BanglaBlend dataset is a comprehensive collection of Bangla (Bengali) sentences meticulously categorized based on two specific forms: Saint(Sadhu) and Common(Cholito). This dataset is comprised of a total 7350 annotated Bangla sentences as well as it is preprocessed dataset where several data preprocessing techniques have been applied. This dataset is designed to facilitate research and development in natural language processing (NLP) and computational linguistics, particularly for Bangla, a widely spoken language in Bangladesh and parts of India. Specially, this dataset can be leveraged for several natural language processing task such as text summarization, text classification, sentiment analysis, automatic language translation.

  18. Runtime comparison.

    • plos.figshare.com
    xls
    Updated Sep 20, 2024
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Shihab Ahmed; Moythry Manir Samia; Maksuda Haider Sayma; Md. Mohsin Kabir; M. F. Mridha (2024). Runtime comparison. [Dataset]. http://doi.org/10.1371/journal.pone.0308050.t011
    Explore at:
    xlsAvailable download formats
    Dataset updated
    Sep 20, 2024
    Dataset provided by
    PLOShttp://plos.org/
    Authors
    Shihab Ahmed; Moythry Manir Samia; Maksuda Haider Sayma; Md. Mohsin Kabir; M. F. Mridha
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    In recent years, the surge in reviews and comments on newspapers and social media has made sentiment analysis a focal point of interest for researchers. Sentiment analysis is also gaining popularity in the Bengali language. However, Aspect-Based Sentiment Analysis is considered a difficult task in the Bengali language due to the shortage of perfectly labeled datasets and the complex variations in the Bengali language. This study used two open-source benchmark datasets of the Bengali language, Cricket, and Restaurant, for our Aspect-Based Sentiment Analysis task. The original work was based on the Random Forest, Support Vector Machine, K-Nearest Neighbors, and Convolutional Neural Network models. In this work, we used the Bidirectional Encoder Representations from Transformers, the Robustly Optimized BERT Approach, and our proposed hybrid transformative Random Forest and Bidirectional Encoder Representations from Transformers (tRF-BERT) models to compare the results with the existing work. After comparing the results, we can clearly see that all the models used in our work achieved better results than any of the previous works on the same dataset. Amongst them, our proposed transformative Random Forest and Bidirectional Encoder Representations from Transformers achieved the highest F1 score and accuracy. The accuracy and F1 score of aspect detection for the Cricket dataset were 0.89 and 0.85, respectively, and for the Restaurant dataset were 0.92 and 0.89 respectively.

  19. BAN-ABSA

    • kaggle.com
    zip
    Updated Oct 2, 2020
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Mahfuz Ahmed Masum (2020). BAN-ABSA [Dataset]. https://www.kaggle.com/datasets/mahfuzahmed/banabsa/suggestions
    Explore at:
    zip(300359 bytes)Available download formats
    Dataset updated
    Oct 2, 2020
    Authors
    Mahfuz Ahmed Masum
    License

    https://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/

    Description

    Dataset

    This dataset was created by Mahfuz Ahmed Masum

    Released under CC0: Public Domain

    Contents

  20. Bangla Financial news articles Dataset

    • kaggle.com
    zip
    Updated Jul 30, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Md. Ashraful Islam (2023). Bangla Financial news articles Dataset [Dataset]. https://www.kaggle.com/mdashrafulislam1998/bangla-financial-news-articles-dataset
    Explore at:
    zip(11501783 bytes)Available download formats
    Dataset updated
    Jul 30, 2023
    Authors
    Md. Ashraful Islam
    License

    Attribution-NonCommercial-ShareAlike 4.0 (CC BY-NC-SA 4.0)https://creativecommons.org/licenses/by-nc-sa/4.0/
    License information was derived automatically

    Description

    Welcome to our Bengali Financial News Sentiment Analysis dataset! This collection comprises 7,695 financial news articles extracted, covering the period from March 3, 2014, to December 29, 2021. Utilizing the powerful web scraping tool "Beautiful Soup 4.4.0" in Python.

    This dataset was a crucial part of our research published in the journal paper titled "Stock Market Prediction of Bangladesh Using Multivariate Long Short-Term Memory with Sentiment Identification." The paper can be accessed and cited at http://doi.org/10.11591/ijece.v13i5.pp5696-5706.

    We are excited to share this unique dataset, which we hope will empower researchers, analysts, and enthusiasts to explore and understand the dynamics of the Bengali financial market through sentiment analysis. Join us on this journey of uncovering the hidden emotions driving market trends and decisions in Bangladesh. Happy analyzing!

Share
FacebookFacebook
TwitterTwitter
Email
Click to copy link
Link copied
Close
Cite
Nuhash Afnan (2020). Bengali Sentiment Dataset [Dataset]. https://www.kaggle.com/nuhashafnan/pseudolabel
Organization logo

Bengali Sentiment Dataset

Explore at:
360 scholarly articles cite this dataset (View in Google Scholar)
zip(631287 bytes)Available download formats
Dataset updated
Jul 25, 2020
Authors
Nuhash Afnan
Description

Dataset

This dataset was created by Nuhash Afnan

Contents

Search
Clear search
Close search
Google apps
Main menu