100+ datasets found
  1. Sentiment Analysis Nepali Dataset

    • kaggle.com
    zip
    Updated Jul 7, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Samir Wagle (2025). Sentiment Analysis Nepali Dataset [Dataset]. https://www.kaggle.com/datasets/sameerwagle/sentimentanalysisnepalidataset
    Explore at:
    zip(2545094027 bytes)Available download formats
    Dataset updated
    Jul 7, 2025
    Authors
    Samir Wagle
    License

    Attribution-NonCommercial 4.0 (CC BY-NC 4.0)https://creativecommons.org/licenses/by-nc/4.0/
    License information was derived automatically

    Description

    web scrapping of comments, extracted available nepali comments existing in huggingface, kaggle. Preprocessed it. Removal of html tag, username, links, emoji, non nepali characters. Lemmatization done. Afrer that clean dataset is with us. That clean dataset is being labeled with LLM as a judge approach using openai o3 model ( premium ) costed us around 20$ for labeling. After that different embedding models were used for embedding. This will act as a benchmark dataset for sentiment analysis of nepali dataset. Please give author a credit if you are using it.

  2. Romanized Nepali Sentiment Analysis Dataset

    • kaggle.com
    zip
    Updated Feb 8, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    erabhash (2026). Romanized Nepali Sentiment Analysis Dataset [Dataset]. https://www.kaggle.com/datasets/erabhash/romanized-nepali-sentiment-analysis-dataset
    Explore at:
    zip(755265 bytes)Available download formats
    Dataset updated
    Feb 8, 2026
    Authors
    erabhash
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    📌 Overview

    This dataset contains annotated Romanized Nepali text samples labeled for sentiment analysis. It is designed to support research in low-resource languages, especially where Romanized forms are widely used on social media platforms such as Facebook, TikTok, and YouTube.

    The dataset was originally prepared for the research work: “Transformer-Based Deep Learning Models for Sentiment Analysis in Romanized Nepali: A Comparative Investigation of BERT and RoBERTa.”

    📁 Dataset Contents Format: CSV

    Columns: - text: Romanized Nepali sentence or comment - label: Sentiment category (Positive, Negative, Neutral)

    ⭐ Key Features - High-quality manual annotations - Focus on low-resource Romanized Nepali text - Balanced sentiment distribution

    Suitable for: -BERT / RoBERTa fine-tuning - Multiclass text classification - Low-resource NLP experiments - Tokenization and preprocessing research

    🎯 Use Cases - Sentiment analysis - Transformer-based model benchmarking - Cross-lingual and low-resource language studies - Data augmentation experiments - Social media text analytics

    ✅ Citation

    If you use this dataset in your research, please cite:

    Abhash Pradhananga, AP. (2024). Transformer-Based Deep Learning Models for Sentiment Analysis in Romanized Nepali: A Comparative Investigation of BERT and RoBERTa.

    Or simply include:

    Dataset provided by Abhash Pradhananga. Please cite the original research paper when using this data.

  3. h

    nepalitext-language-model-dataset

    • huggingface.co
    Updated Jan 31, 2023
    + more versions
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Utsav Maskey (2023). nepalitext-language-model-dataset [Dataset]. https://huggingface.co/datasets/Sakonii/nepalitext-language-model-dataset
    Explore at:
    CroissantCroissant is a format for machine-learning datasets. Learn more about this at mlcommons.org/croissant.
    Dataset updated
    Jan 31, 2023
    Authors
    Utsav Maskey
    License

    https://choosealicense.com/licenses/cc0-1.0/https://choosealicense.com/licenses/cc0-1.0/

    Description

    Dataset Card for "nepalitext-language-model-dataset"

      Dataset Summary
    

    "NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia.

      Supported Tasks and Leaderboards
    

    This dataset is intended to pre-train language models and word representations on Nepali Language.

      Languages
    

    The data is… See the full description on the dataset page: https://huggingface.co/datasets/Sakonii/nepalitext-language-model-dataset.

  4. h

    openSLR-Nepali

    • huggingface.co
    Updated Nov 22, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Aananda Giri (2025). openSLR-Nepali [Dataset]. https://huggingface.co/datasets/Aananda-giri/openSLR-Nepali
    Explore at:
    Dataset updated
    Nov 22, 2025
    Authors
    Aananda Giri
    License

    Attribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
    License information was derived automatically

    Description

    OpenSLR Nepali Speech Dataset (Preprocessed)

      Dataset Description
    

    This is a preprocessed version of the Nepali speech dataset from OpenSLR, ready for training speech models including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS).

      Dataset Statistics
    

    Total Audio Files: 118,231 Total Duration: 57.34 hours Sample Rate: 16kHz Channels: Mono Format: WAV

      Preprocessing Applied
    

    Text Preprocessing:

    Text cleaning and… See the full description on the dataset page: https://huggingface.co/datasets/Aananda-giri/openSLR-Nepali.

  5. h

    English-Hindi-Nepali-Retrieval

    • huggingface.co
    Updated Jul 13, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    nish pd (2026). English-Hindi-Nepali-Retrieval [Dataset]. https://huggingface.co/datasets/nishpish/English-Hindi-Nepali-Retrieval
    Explore at:
    Dataset updated
    Jul 13, 2026
    Authors
    nish pd
    License

    MIT Licensehttps://opensource.org/licenses/MIT
    License information was derived automatically

    Area covered
    Nepal
    Description

    nishpish/English-Hindi-Nepali-Retrieval dataset hosted on Hugging Face and contributed by the HF Datasets community

  6. Nepali Handwritten Images for text detection

    • kaggle.com
    zip
    Updated Sep 23, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Sweekar Dahal (2023). Nepali Handwritten Images for text detection [Dataset]. https://www.kaggle.com/datasets/sweekardahal/nepali-handwritten-images-for-text-detection
    Explore at:
    zip(1304832759 bytes)Available download formats
    Dataset updated
    Sep 23, 2023
    Authors
    Sweekar Dahal
    License

    https://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/

    Description

    Implementations:

    1. https://github.com/R4j4n/Nepali-Text-Detection-DBnet
    2. https://github.com/dahalsweekar/TextBoxes---Compatible-with-python-3.10 (outdated)

    Description:

    We present the Nepali Handwriting Dataset (NHD), which is a collection of camera-captured images of Nepali handwritten text from various regions in Nepal. The dataset aims to provide a benchmark for researchers to explore new techniques in handwriting detection and recognition. We also present benchmark results for text localization and recognition using well-established deep-learning frameworks. The dataset and benchmark results are available here.

    Key Features:

    The role of data collection and preprocessing in the research on handwritten text detection cannot be overstated. It is a crucial aspect that plays a significant role in obtaining a comprehensive and diverse dataset. To this end, the researchers personally collected 1,000 mobile phone-captured data samples from various sources, including schools, government offices, universities, and student councils.

    The dataset was carefully curated to encompass three distinct categories based on age groups, namely kids, youth, and adults, with 599, 152, and 249 samples, respectively. Each of the 1,000 pages was meticulously annotated by the researchers to ensure accurate labeling and create a reliable dataset. The data collection process focused on capturing a wide range of handwriting styles and variations prevalent among different age groups and settings.

    The collected dataset served as a valuable resource for training and evaluating the handwritten text detection models in the research. It provided a rich and diverse set of data that enabled the researchers to develop robust models capable of accurately detecting handwritten text across different age groups and settings.

    Use Cases:

    1. Real-Time Text Detection
    2. Text Recognition

    Results:

    You can find its implementation here: https://github.com/R4j4n/Nepali-Text-Detection-DBnet

    Recall: 0.9069154470416869

    Precision: 0.9178659178659179

    HMean: 0.9123578206927347

    Test Image:

    https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4786384%2Ff8d9aa282a42848b359aeeb021b97937%2Foutput.png?generation=1695433752833462&alt=media" alt="">

    If you find this dataset useful, your support through an upvote would be greatly appreciated ❤️🙂

    Thank you

  7. R

    Nepali Characters Dataset

    • universe.roboflow.com
    zip
    Updated Feb 2, 2022
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Sharad Pathak (2022). Nepali Characters Dataset [Dataset]. https://universe.roboflow.com/sharad-pathak/nepali-dataset-characters/dataset/2
    Explore at:
    zipAvailable download formats
    Dataset updated
    Feb 2, 2022
    Dataset authored and provided by
    Sharad Pathak
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Variables measured
    Text Bounding Boxes
    Description

    Nepali Dataset Characters

    ## Overview
    
    Nepali Dataset Characters is a dataset for object detection tasks - it contains Text annotations for 275 images.
    
    ## Getting Started
    
    You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
    
      ## License
    
      This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
    
  8. p

    Nepali Datasets for AI Training

    • pangeanic.com
    text
    Updated May 19, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Pangeanic (2026). Nepali Datasets for AI Training [Dataset]. https://pangeanic.com/nepali-datasets-for-ai-training
    Explore at:
    textAvailable download formats
    Dataset updated
    May 19, 2026
    Dataset provided by
    Pangeanic S.L
    Authors
    Pangeanic
    License

    https://pangeanic.com/contact-ushttps://pangeanic.com/contact-us

    Area covered
    Nepal, Biratnagar, Pokhara, Kathmandu, South Asia, Nepaliland
    Variables measured
    Nepali OCR text, Retail terminology, Fintech terminology, Conversational Nepali, Enterprise communication, Social media communication, Nepali speech transcription, Customer support interactions, Digital commerce interactions, Nepali-English code-switching, and 2 more
    Measurement technique
    Multilingual dataset curation, Quality control sampling, Speech transcription, Data cleaning, Expert review, Human-in-the-loop annotation, Linguistic quality assurance, Metadata enrichment
    Description

    Enterprise-grade Nepali datasets covering conversational Nepali, Nepali-English code-switching, customer communication, OCR, speech, text, enterprise NLP and South Asian multilingual AI workflows.

  9. MSVD Nepali Dataset

    • kaggle.com
    zip
    Updated Jul 30, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Bipesh Raj Subedi (2023). MSVD Nepali Dataset [Dataset]. https://www.kaggle.com/datasets/bipeshrajsubedi/msvd-nepali-dataset
    Explore at:
    zip(1843221738 bytes)Available download formats
    Dataset updated
    Jul 30, 2023
    Authors
    Bipesh Raj Subedi
    Description

    The dataset contains translated Nepali captions along with its original text in a csv file. The dataset can be used for educational purpose with proper citation. The dataset is in raw format and needs to be preprocessed as per your application.

    Credit: The original dataset was collected from: @InProceedings{chen:acl11, title = "Collecting Highly Parallel Data for Paraphrase Evaluation", author = "David L. Chen and William B. Dolan", booktitle = "Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics (ACL-2011)", address = "Portland, OR", month = "June", year = 2011 } Link to original data: https://www.cs.utexas.edu/users/ml/clamp/videoDescription/

  10. Nepali-News-summary

    • kaggle.com
    zip
    Updated Jun 9, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Adhikary Kishan (2025). Nepali-News-summary [Dataset]. https://www.kaggle.com/datasets/adhikarykishan/nepali-news-summary
    Explore at:
    zip(60115664 bytes)Available download formats
    Dataset updated
    Jun 9, 2025
    Authors
    Adhikary Kishan
    License

    MIT Licensehttps://opensource.org/licenses/MIT
    License information was derived automatically

    Area covered
    नेपाल
    Description

    📰 Nepali News Summary Dataset (CSV Format) This dataset contains 51k pairs of Nepali news articles and their corresponding summaries in CSV format. It is designed to support natural language processing tasks such as text summarization in the Nepali language.

    📁 Dataset Format The dataset is provided in a .csv file with the following structure:

    Column Name Description article Full news article in Nepali summary Corresponding summary in Nepali

    📊 Dataset Overview Total Samples: 51000

    Manually created summaries: 5,000

    Synthetic summaries (via Gemini API): 44,000

    Even though there's no source column, the dataset was built from diverse Nepali news portals, with this approximate distribution:

    Source Approx. Count OnlineKhabar 20,000 (4,000 manually scraped) Setopati 7,000 Baahrakhari 1,200 MyRepublica 2,000 Karobar 10,000 Lokpath 6,000 Ratopati 4,000

    Note: This breakdown is just for informational purposes — it's not included in the dataset.

    ✅ Example Row (CSV) csv Copy Edit article,summary "प्रधानमन्त्रीले आज नयाँ नीति घोषणा गरे...", "प्रधानमन्त्रीले नयाँ नीति सार्वजनिक गरे।"

    💡 Use Cases Training and benchmarking summarization models for low-resource languages

    Fine-tuning LLMs and encoder-decoder models for Nepali

    Research in machine translation, text generation, and NLP for South Asian languages

  11. h

    Nepali

    • huggingface.co
    Updated Feb 8, 2024
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Sunil Chaudhary (2024). Nepali [Dataset]. https://huggingface.co/datasets/SunilC/Nepali
    Explore at:
    Dataset updated
    Feb 8, 2024
    Authors
    Sunil Chaudhary
    License

    MIT Licensehttps://opensource.org/licenses/MIT
    License information was derived automatically

    Description

    SunilC/Nepali dataset hosted on Hugging Face and contributed by the HF Datasets community

  12. h

    English-Nepali-Translation-Dataset

    • huggingface.co
    Updated Jan 31, 2024
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Ashok Poudel (2024). English-Nepali-Translation-Dataset [Dataset]. https://huggingface.co/datasets/ashokpoudel/English-Nepali-Translation-Dataset
    Explore at:
    Dataset updated
    Jan 31, 2024
    Authors
    Ashok Poudel
    Description

    ashokpoudel/English-Nepali-Translation-Dataset dataset hosted on Hugging Face and contributed by the HF Datasets community

  13. OpenSLR Large Nepali dataset cleaned

    • kaggle.com
    zip
    Updated Mar 10, 2024
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Aashish Kumar Sah (2024). OpenSLR Large Nepali dataset cleaned [Dataset]. https://www.kaggle.com/datasets/aashish12321/openslr-large-nepali-dataset-cleaned/code
    Explore at:
    zip(4035232018 bytes)Available download formats
    Dataset updated
    Mar 10, 2024
    Authors
    Aashish Kumar Sah
    Description

    Dataset

    This dataset was created by Aashish Kumar Sah

    Contents

  14. h

    slr-combined-nepali-tts

    • huggingface.co
    Updated Jul 12, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Kiran Silwal2 (2026). slr-combined-nepali-tts [Dataset]. https://huggingface.co/datasets/lilgoose7777/slr-combined-nepali-tts
    Explore at:
    Dataset updated
    Jul 12, 2026
    Authors
    Kiran Silwal2
    Description

    lilgoose7777/slr-combined-nepali-tts dataset hosted on Hugging Face and contributed by the HF Datasets community

  15. Nepali Fake News Detection Dataset

    • kaggle.com
    zip
    Updated Jan 27, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Ashok Nepal (2026). Nepali Fake News Detection Dataset [Dataset]. https://www.kaggle.com/datasets/ashoknepal/nepali-fake-news-detection/discussion
    Explore at:
    zip(3985706 bytes)Available download formats
    Dataset updated
    Jan 27, 2026
    Authors
    Ashok Nepal
    License

    Attribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
    License information was derived automatically

    Description

    The Nepali Fake News Detection Dataset is a curated research dataset designed to support fake news detection, misinformation analysis, and Natural Language Processing (NLP) research for the Nepali language.

    The dataset contains Nepali-language news content collected from publicly available social media platforms and official online news portals. Each data instance has been reviewed and labeled to distinguish between real and fake news, making the dataset suitable for supervised machine learning tasks.

    This dataset aims to address the lack of high-quality labeled Nepali datasets and can be used by researchers, students, and practitioners for academic research, benchmarking, and experimental analysis.

  16. h

    Nepali

    • huggingface.co
    Updated Sep 12, 2024
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Bibek (2024). Nepali [Dataset]. https://huggingface.co/datasets/ibibek/Nepali
    Explore at:
    Dataset updated
    Sep 12, 2024
    Authors
    Bibek
    Description

    ibibek/Nepali dataset hosted on Hugging Face and contributed by the HF Datasets community

  17. r

    Nepali language — United States

    • rascasse.com
    html, json
    Updated May 25, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Rascasse (2026). Nepali language — United States [Dataset]. https://rascasse.com/explore/us/nepali-language-69240
    Explore at:
    json, htmlAvailable download formats
    Dataset updated
    May 25, 2026
    Dataset authored and provided by
    Rascasse
    License

    https://rascasse.com/terms/https://rascasse.com/terms/

    Time period covered
    2026
    Area covered
    United States
    Variables measured
    Male share, Average age, Female share, Audience size
    Measurement technique
    Search-behavior signal aggregation
    Description

    Nepali language audience profile for United States.

  18. r

    Nepali language Audience Profile — Germany 2026

    • rascasse.com
    Updated May 12, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Rascasse (2026). Nepali language Audience Profile — Germany 2026 [Dataset]. https://rascasse.com/explore/de/nepali-language-69240
    Explore at:
    Dataset updated
    May 12, 2026
    Dataset authored and provided by
    Rascasse
    License

    https://rascasse.com/termshttps://rascasse.com/terms

    Area covered
    Germany
    Variables measured
    Fan Count, Median Age, Female Share
    Description

    Demographic, psychographic, geographic and brand-affinity data for the Nepali language audience in Germany, sourced from Rascasse's panel of 12+ social and digital signals.

  19. h

    XLSum-nepali-summerization-dataset

    • huggingface.co
    Updated Mar 6, 2024
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Sanjeev Bhandari (2024). XLSum-nepali-summerization-dataset [Dataset]. https://huggingface.co/datasets/realsanjeev/XLSum-nepali-summerization-dataset
    Explore at:
    Dataset updated
    Mar 6, 2024
    Authors
    Sanjeev Bhandari
    License

    MIT Licensehttps://opensource.org/licenses/MIT
    License information was derived automatically

    Description

    realsanjeev/XLSum-nepali-summerization-dataset dataset hosted on Hugging Face and contributed by the HF Datasets community

  20. r

    Nepali language Audience Profile — India 2026

    • rascasse.com
    Updated May 12, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Rascasse (2026). Nepali language Audience Profile — India 2026 [Dataset]. https://rascasse.com/explore/in/nepali-language-69240
    Explore at:
    Dataset updated
    May 12, 2026
    Dataset authored and provided by
    Rascasse
    License

    https://rascasse.com/termshttps://rascasse.com/terms

    Area covered
    India
    Variables measured
    Fan Count, Median Age, Female Share
    Description

    Demographic, psychographic, geographic and brand-affinity data for the Nepali language audience in India, sourced from Rascasse's panel of 12+ social and digital signals.

Share
FacebookFacebook
TwitterTwitter
Email
Click to copy link
Link copied
Close
Cite
Samir Wagle (2025). Sentiment Analysis Nepali Dataset [Dataset]. https://www.kaggle.com/datasets/sameerwagle/sentimentanalysisnepalidataset
Organization logo

Sentiment Analysis Nepali Dataset

100k dataset with labeled, embeded using different pretrained model.-

Explore at:
6 scholarly articles cite this dataset (View in Google Scholar)
zip(2545094027 bytes)Available download formats
Dataset updated
Jul 7, 2025
Authors
Samir Wagle
License

Attribution-NonCommercial 4.0 (CC BY-NC 4.0)https://creativecommons.org/licenses/by-nc/4.0/
License information was derived automatically

Description

web scrapping of comments, extracted available nepali comments existing in huggingface, kaggle. Preprocessed it. Removal of html tag, username, links, emoji, non nepali characters. Lemmatization done. Afrer that clean dataset is with us. That clean dataset is being labeled with LLM as a judge approach using openai o3 model ( premium ) costed us around 20$ for labeling. After that different embedding models were used for embedding. This will act as a benchmark dataset for sentiment analysis of nepali dataset. Please give author a credit if you are using it.

Search
Clear search
Close search
Google apps
Main menu