2 datasets found
  1. m

    BanglaSER: A Bangla speech emotion recognition dataset

    • data.mendeley.com
    Updated Mar 14, 2022
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Rakesh Kumar Das (2022). BanglaSER: A Bangla speech emotion recognition dataset [Dataset]. http://doi.org/10.17632/t9h6p943xy.5
    Explore at:
    Dataset updated
    Mar 14, 2022
    Authors
    Rakesh Kumar Das
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    BanglaSER is a Bangla language-based speech emotion recognition dataset. It consists of speech-audio data of 34 participating speakers from diverse age groups between 19 and 47 years, with a balanced 17 male and 17 female nonprofessional participating actors. This dataset contains 1467 Bangla speech-audio recordings of five rudimentary human emotional states, namely angry, happy, neutral, sad, and surprise. Three trials are conducted for each emotional state. Hence, the total number of recordings involves 3 statements × 3 repetitions × 4 emotional states (angry, happy, sad, and surprise) × 34 participating speakers = 1224 recordings + 3 statements × 3 repetitions × 1 emotional state (neutral) × 27 participating speakers = 243 recordings, making the total number of recordings of 1467. BanglaSER dataset is collected by recording through smartphones, and laptops, having a balanced number of recordings in each category with evenly distributed participating male and female actors, preserves the real-life environment, and would serve as an essential training dataset for the speech emotion recognition model in terms of generalization. BanglaSER is compatible with various deep learning architectures such as CNN, LSTM, BiLSTM etc.

  2. BanglaSER: Bangla Audio for Emotion Recognition

    • kaggle.com
    Updated Aug 27, 2024
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Evil Spirit05 (2024). BanglaSER: Bangla Audio for Emotion Recognition [Dataset]. https://www.kaggle.com/datasets/evilspirit05/emotion
    Explore at:
    CroissantCroissant is a format for machine-learning datasets. Learn more about this at mlcommons.org/croissant.
    Dataset updated
    Aug 27, 2024
    Dataset provided by
    Kaggle
    Authors
    Evil Spirit05
    License

    MIT Licensehttps://opensource.org/licenses/MIT
    License information was derived automatically

    Description
    BanglaSER is a specialized dataset designed for the task of Bangla speech emotion recognition. This dataset includes a rich collection of speech-audio recordings that capture a variety of fundamental human emotions. It is curated to support research and development in the field of speech emotion recognition, particularly for the Bangla language, and is suitable for various deep learning architectures.
    

    Dataset Composition:

    • Total Number of Recordings: 1,467
    • Number of Speakers: 34 (17 male and 17 female)
    • Age Range of Speakers: 19 to 47 years
    • Recording Devices: Smartphones and laptops

    Emotional States Covered:

    • Angry
    • Happy
    • Neutral
    • Sad
    • Surprise

    Recording Structure:

    Each emotional state is represented by:

    • 3 Statements spoken three times by each participant.
    • For Angry, Happy, Sad, and Surprise: 3 statements × 3 repetitions × 34 speakers = 1,224 recordings.
    • For Neutral: 3 statements × 3 repetitions × 27 speakers = 243 recordings

    Key Features:

    Balanced Representation:

    • The dataset is carefully balanced with an equal number of male and female participants, ensuring that the recordings reflect diverse voices and emotional expressions.
    • Emotions are evenly distributed across the dataset, providing a robust basis for training and evaluating emotion recognition models.

    Realistic Recording Conditions:

    • Recordings are made using commonly available devices, such as smartphones and laptops, which helps in preserving the naturalistic quality of the audio.
    • The dataset reflects real-life acoustic environments, making it more applicable to real-world applications.

    Deep Learning Compatibility:

    • BanglaSER is designed to be compatible with various deep learning architectures, including Convolutional Neural Networks (CNNs), Long Short-Term Memory Networks (LSTMs), and Bidirectional LSTMs (BiLSTMs).
    • The dataset can be used for a range of tasks, from emotion classification to sentiment analysis, and more.

    Usage and Applications:

    • Emotion Recognition Models: BanglaSER provides a diverse set of recordings that are ideal for training models to recognize and classify emotions in Bangla speech.
    • Benchmarking and Evaluation: The dataset serves as a benchmark for evaluating the performance of emotion recognition systems and can help in comparing different model architectures and techniques.
    • Research and Development: Researchers can use BanglaSER to explore new methods in speech emotion recognition, develop novel algorithms, and enhance the understanding of emotion in speech.

    Dataset Access:

    Download Link: https://data.mendeley.com/datasets/t9h6p943xy/5

    • Documentation: Detailed documentation and guidelines for using the dataset are provided to assist users in effectively leveraging the data.

    Acknowledgments:

    We extend our gratitude to the contributors and participants who made this dataset possible. Their efforts have greatly enriched the field of speech emotion recognition and provided valuable resources for the community.
    
    Feel free to explore the dataset and utilize it in your research and projects. We look forward to seeing the innovative applications and advancements that will emerge from the use of BanglaSER
    
  3. Not seeing a result you expected?
    Learn how you can add new datasets to our index.

Share
FacebookFacebook
TwitterTwitter
Email
Click to copy link
Link copied
Close
Cite
Rakesh Kumar Das (2022). BanglaSER: A Bangla speech emotion recognition dataset [Dataset]. http://doi.org/10.17632/t9h6p943xy.5

BanglaSER: A Bangla speech emotion recognition dataset

Explore at:
Dataset updated
Mar 14, 2022
Authors
Rakesh Kumar Das
License

Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically

Description

BanglaSER is a Bangla language-based speech emotion recognition dataset. It consists of speech-audio data of 34 participating speakers from diverse age groups between 19 and 47 years, with a balanced 17 male and 17 female nonprofessional participating actors. This dataset contains 1467 Bangla speech-audio recordings of five rudimentary human emotional states, namely angry, happy, neutral, sad, and surprise. Three trials are conducted for each emotional state. Hence, the total number of recordings involves 3 statements × 3 repetitions × 4 emotional states (angry, happy, sad, and surprise) × 34 participating speakers = 1224 recordings + 3 statements × 3 repetitions × 1 emotional state (neutral) × 27 participating speakers = 243 recordings, making the total number of recordings of 1467. BanglaSER dataset is collected by recording through smartphones, and laptops, having a balanced number of recordings in each category with evenly distributed participating male and female actors, preserves the real-life environment, and would serve as an essential training dataset for the speech emotion recognition model in terms of generalization. BanglaSER is compatible with various deep learning architectures such as CNN, LSTM, BiLSTM etc.

Search
Clear search
Close search
Google apps
Main menu