Facebook
TwitterAttribution-NonCommercial 4.0 (CC BY-NC 4.0)https://creativecommons.org/licenses/by-nc/4.0/
License information was derived automatically
web scrapping of comments, extracted available nepali comments existing in huggingface, kaggle. Preprocessed it. Removal of html tag, username, links, emoji, non nepali characters. Lemmatization done. Afrer that clean dataset is with us. That clean dataset is being labeled with LLM as a judge approach using openai o3 model ( premium ) costed us around 20$ for labeling. After that different embedding models were used for embedding. This will act as a benchmark dataset for sentiment analysis of nepali dataset. Please give author a credit if you are using it.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
📌 Overview
This dataset contains annotated Romanized Nepali text samples labeled for sentiment analysis. It is designed to support research in low-resource languages, especially where Romanized forms are widely used on social media platforms such as Facebook, TikTok, and YouTube.
The dataset was originally prepared for the research work: “Transformer-Based Deep Learning Models for Sentiment Analysis in Romanized Nepali: A Comparative Investigation of BERT and RoBERTa.”
📁 Dataset Contents Format: CSV
Columns: - text: Romanized Nepali sentence or comment - label: Sentiment category (Positive, Negative, Neutral)
⭐ Key Features - High-quality manual annotations - Focus on low-resource Romanized Nepali text - Balanced sentiment distribution
Suitable for: -BERT / RoBERTa fine-tuning - Multiclass text classification - Low-resource NLP experiments - Tokenization and preprocessing research
🎯 Use Cases - Sentiment analysis - Transformer-based model benchmarking - Cross-lingual and low-resource language studies - Data augmentation experiments - Social media text analytics
✅ Citation
If you use this dataset in your research, please cite:
Abhash Pradhananga, AP. (2024). Transformer-Based Deep Learning Models for Sentiment Analysis in Romanized Nepali: A Comparative Investigation of BERT and RoBERTa.
Or simply include:
Dataset provided by Abhash Pradhananga. Please cite the original research paper when using this data.
Facebook
Twitterhttps://choosealicense.com/licenses/cc0-1.0/https://choosealicense.com/licenses/cc0-1.0/
Dataset Card for "nepalitext-language-model-dataset"
Dataset Summary
"NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia.
Supported Tasks and Leaderboards
This dataset is intended to pre-train language models and word representations on Nepali Language.
Languages
The data is… See the full description on the dataset page: https://huggingface.co/datasets/Sakonii/nepalitext-language-model-dataset.
Facebook
TwitterAttribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
License information was derived automatically
OpenSLR Nepali Speech Dataset (Preprocessed)
Dataset Description
This is a preprocessed version of the Nepali speech dataset from OpenSLR, ready for training speech models including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS).
Dataset Statistics
Total Audio Files: 118,231 Total Duration: 57.34 hours Sample Rate: 16kHz Channels: Mono Format: WAV
Preprocessing Applied
Text Preprocessing:
Text cleaning and… See the full description on the dataset page: https://huggingface.co/datasets/Aananda-giri/openSLR-Nepali.
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
nishpish/English-Hindi-Nepali-Retrieval dataset hosted on Hugging Face and contributed by the HF Datasets community
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
Implementations:
Description:
We present the Nepali Handwriting Dataset (NHD), which is a collection of camera-captured images of Nepali handwritten text from various regions in Nepal. The dataset aims to provide a benchmark for researchers to explore new techniques in handwriting detection and recognition. We also present benchmark results for text localization and recognition using well-established deep-learning frameworks. The dataset and benchmark results are available here.
Key Features:
The role of data collection and preprocessing in the research on handwritten text detection cannot be overstated. It is a crucial aspect that plays a significant role in obtaining a comprehensive and diverse dataset. To this end, the researchers personally collected 1,000 mobile phone-captured data samples from various sources, including schools, government offices, universities, and student councils.
The dataset was carefully curated to encompass three distinct categories based on age groups, namely kids, youth, and adults, with 599, 152, and 249 samples, respectively. Each of the 1,000 pages was meticulously annotated by the researchers to ensure accurate labeling and create a reliable dataset. The data collection process focused on capturing a wide range of handwriting styles and variations prevalent among different age groups and settings.
The collected dataset served as a valuable resource for training and evaluating the handwritten text detection models in the research. It provided a rich and diverse set of data that enabled the researchers to develop robust models capable of accurately detecting handwritten text across different age groups and settings.
Use Cases:
Results:
You can find its implementation here: https://github.com/R4j4n/Nepali-Text-Detection-DBnet
Recall: 0.9069154470416869
Precision: 0.9178659178659179
HMean: 0.9123578206927347
Test Image:
https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4786384%2Ff8d9aa282a42848b359aeeb021b97937%2Foutput.png?generation=1695433752833462&alt=media" alt="">
If you find this dataset useful, your support through an upvote would be greatly appreciated ❤️🙂
Thank you
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
## Overview
Nepali Dataset Characters is a dataset for object detection tasks - it contains Text annotations for 275 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
Facebook
Twitterhttps://pangeanic.com/contact-ushttps://pangeanic.com/contact-us
Enterprise-grade Nepali datasets covering conversational Nepali, Nepali-English code-switching, customer communication, OCR, speech, text, enterprise NLP and South Asian multilingual AI workflows.
Facebook
TwitterThe dataset contains translated Nepali captions along with its original text in a csv file. The dataset can be used for educational purpose with proper citation. The dataset is in raw format and needs to be preprocessed as per your application.
Credit: The original dataset was collected from: @InProceedings{chen:acl11, title = "Collecting Highly Parallel Data for Paraphrase Evaluation", author = "David L. Chen and William B. Dolan", booktitle = "Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics (ACL-2011)", address = "Portland, OR", month = "June", year = 2011 } Link to original data: https://www.cs.utexas.edu/users/ml/clamp/videoDescription/
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
📰 Nepali News Summary Dataset (CSV Format) This dataset contains 51k pairs of Nepali news articles and their corresponding summaries in CSV format. It is designed to support natural language processing tasks such as text summarization in the Nepali language.
📁 Dataset Format The dataset is provided in a .csv file with the following structure:
Column Name Description article Full news article in Nepali summary Corresponding summary in Nepali
📊 Dataset Overview Total Samples: 51000
Manually created summaries: 5,000
Synthetic summaries (via Gemini API): 44,000
Even though there's no source column, the dataset was built from diverse Nepali news portals, with this approximate distribution:
Source Approx. Count OnlineKhabar 20,000 (4,000 manually scraped) Setopati 7,000 Baahrakhari 1,200 MyRepublica 2,000 Karobar 10,000 Lokpath 6,000 Ratopati 4,000
Note: This breakdown is just for informational purposes — it's not included in the dataset.
✅ Example Row (CSV) csv Copy Edit article,summary "प्रधानमन्त्रीले आज नयाँ नीति घोषणा गरे...", "प्रधानमन्त्रीले नयाँ नीति सार्वजनिक गरे।"
💡 Use Cases Training and benchmarking summarization models for low-resource languages
Fine-tuning LLMs and encoder-decoder models for Nepali
Research in machine translation, text generation, and NLP for South Asian languages
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
SunilC/Nepali dataset hosted on Hugging Face and contributed by the HF Datasets community
Facebook
Twitterashokpoudel/English-Nepali-Translation-Dataset dataset hosted on Hugging Face and contributed by the HF Datasets community
Facebook
TwitterThis dataset was created by Aashish Kumar Sah
Facebook
Twitterlilgoose7777/slr-combined-nepali-tts dataset hosted on Hugging Face and contributed by the HF Datasets community
Facebook
TwitterAttribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
License information was derived automatically
The Nepali Fake News Detection Dataset is a curated research dataset designed to support fake news detection, misinformation analysis, and Natural Language Processing (NLP) research for the Nepali language.
The dataset contains Nepali-language news content collected from publicly available social media platforms and official online news portals. Each data instance has been reviewed and labeled to distinguish between real and fake news, making the dataset suitable for supervised machine learning tasks.
This dataset aims to address the lack of high-quality labeled Nepali datasets and can be used by researchers, students, and practitioners for academic research, benchmarking, and experimental analysis.
Facebook
Twitterhttps://rascasse.com/terms/https://rascasse.com/terms/
Nepali language audience profile for United States.
Facebook
Twitterhttps://rascasse.com/termshttps://rascasse.com/terms
Demographic, psychographic, geographic and brand-affinity data for the Nepali language audience in Germany, sourced from Rascasse's panel of 12+ social and digital signals.
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
realsanjeev/XLSum-nepali-summerization-dataset dataset hosted on Hugging Face and contributed by the HF Datasets community
Facebook
Twitterhttps://rascasse.com/termshttps://rascasse.com/terms
Demographic, psychographic, geographic and brand-affinity data for the Nepali language audience in India, sourced from Rascasse's panel of 12+ social and digital signals.
Facebook
TwitterAttribution-NonCommercial 4.0 (CC BY-NC 4.0)https://creativecommons.org/licenses/by-nc/4.0/
License information was derived automatically
web scrapping of comments, extracted available nepali comments existing in huggingface, kaggle. Preprocessed it. Removal of html tag, username, links, emoji, non nepali characters. Lemmatization done. Afrer that clean dataset is with us. That clean dataset is being labeled with LLM as a judge approach using openai o3 model ( premium ) costed us around 20$ for labeling. After that different embedding models were used for embedding. This will act as a benchmark dataset for sentiment analysis of nepali dataset. Please give author a credit if you are using it.