Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
Welcome to an exceptional dataset meticulously crafted for training state-of-the-art language models such as Gemma, Llama 2, Orca, and more.
Dataset Highlights - Challenging Questions : Immerse your language models in various Python programming questions designed to stimulate cognitive growth. - Real-world Inputs : Provide your models with authentic input scenarios, ensuring they are well-equipped to handle practical coding challenges. - Accurate Answers : Sharpen the precision of your language models by exposing them to meticulously crafted Python code solutions.
How to Get Started - Download : Grab a copy of the dataset and inject new life into your language models. - Build Brilliance : Watch your LLMs evolve as they engage with the challenging questions and nuanced coding scenarios. - Share & Collaborate : Join the Kaggle community to discuss, share insights, and collaborate with fellow enthusiasts.
Unleash the full potential of your language models with this dataset. Elevate your LLM training experience and witness unprecedented growth in language understanding and coding prowess. Happy coding !
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
A common question for those new and familiar to computer science and software engineering is what is the most best and/or most popular programming language. It is very difficult to give a definitive answer, as there are a seemingly indefinite number of metrics that can define the 'best' or 'most popular' programming language.
One such metric that can be used to define a 'popular' programming language is the number of projects and files that are made using that programming language. As GitHub is the most popular public collaboration and file-sharing platform, analyzing the languages that are used for repositories, PRs, and issues on GitHub and be a good indicator for the popularity of a language.
This dataset contains statistics about the programming languages used for repositories, PRs, and issues on GitHub. The data is from 2011 to 2021.
This data was queried and aggregated from BigQuery's public github_repos and githubarchive datasets.
Only data for public GitHub repositories, and their corresponding PRs/issues, have their data available publicly. Thus, this dataset is only based on public repositories, which may not be fully representative of all repositories on GitHub.
Facebook
Twitterhttps://choosealicense.com/licenses/other/https://choosealicense.com/licenses/other/
Dataset Card for The Stack
Changelog
Release Description
v1.0 Initial release of the Stack. Included 30 programming languages and 18 permissive licenses. Note: Three included licenses (MPL/EPL/LGPL) are considered weak copyleft licenses. The resulting near-deduplicated dataset is 3TB in size.
v1.1 The three copyleft licenses ((MPL/EPL/LGPL) were excluded and the list of permissive licenses extended to 193 licenses in total. The list of programming… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack.
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
🧠ProgrammingDataset
A high-quality, production-grade dataset of programming code snippets across multiple languages, collected and curated manually to support research in code generation, analysis, and educational tools.
📌 Dataset Summary
Field Description
Rows 100+ code samples
Languages Python, JavaScript, C++, Java, etc.
Tasks Data structures, algorithms, system utilities
Format Excel (.xlsx) and CSV
License MIT
Each entry includes:
id:… See the full description on the dataset page: https://huggingface.co/datasets/kaiiddo/ProgrammingDataset.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
Dataset Description:
Nemotron-Competitive-Programming-v1 is a large-scale synthetic coding and reasoning dataset designed to push LLM performance on challenging programming and systems tasks. It combines Python and C++ samples across unique competitive programming questions.
Beyond problem solving, the dataset includes InfiniByte, a cross-domain subset with problems derived from scientific fields.
This dataset is ready for commercial use.
Competitive Coding
The… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Competitive-Programming-v1.
Facebook
TwitterAs of 2025, JavaScript and HTML/CSS are the most commonly used programming languages among software developers around the world, with more than 66 percent of respondents stating that they used JavaScript and just around 61.9 percent using HTML/CSS. Python, SQL, and Bash/Shell rounded out the top five most widely used programming languages around the world. Programming languages At a very basic level, programming languages serve as sets of instructions that direct computers on how to behave and carry out tasks. Thanks to the increased prevalence of, and reliance on, computers and electronic devices in today’s society, these languages play a crucial role in the everyday lives of people around the world. An increasing number of people are interested in furthering their understanding of these tools through courses and bootcamps, while current developers are constantly seeking new languages and resources to learn to add to their skills. Furthermore, programming knowledge is becoming an important skill to possess within various industries throughout the business world. Job seekers with skills in Python, R, and SQL will find their knowledge to be among the most highly desirable data science skills and likely assist in their search for employment.
Facebook
TwitterMulti-Round Programming Conversations
Based on previous evol-codealpaca-v1 dataset with added sampled questions from stackoverflow, crossvalidated and make it multiround! It should be more suited to train a code assistant which works side by side.
Tasks included in here:
Data science, statistic, programming questions
Code translation : translate a short function from Python, Golang, C++, Java, Javascript
Code fixing : Fix randomly corrupts characters with no tab… See the full description on the dataset page: https://huggingface.co/datasets/theblackcat102/multiround-programming-convo.
Facebook
TwitterThis dataset is a curated collection of competitive programming problem statements and their structured specifications, built to support research and experimentation in problem understanding, input–output extraction, and dataset creation for machine learning and large language models.
The dataset was originally created by scraping problem statements from online competitive programming platforms during college, and later updated and cleaned after platform changes such as restricted access for paid users and website restructuring. The goal was to preserve high-quality raw problem text along with clearly separated and standardized Input and Output specifications.
🔹 How the Dataset Was Built
After scraping, the raw HTML/text data was:
Input and Output sections were extracted and rewritten into a structured, machine-readable format, preserving the exact meaning defined in the original problem statements.
During later re-runs of the scraping pipeline, the dataset accounts for:
This makes the dataset suitable for studying robustness in data collection pipelines and long-term dataset maintenance.
🔹 What the Dataset Contains
Each entry in the dataset typically includes a combination of:
(Optional, depending on version)
The dataset focuses strictly on problem descriptions and specifications, without including solutions, algorithms, or editorial logic.
🔹 Intended Use Cases
This dataset is suitable for:
It is particularly useful for projects involving:
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
By Tarun Bisht (From Huggingface) [source]
The python_code_instructions_18k_alpaca dataset is a comprehensive training dataset specifically curated for researchers and developers involved in the analysis and comprehension of Python code instructions. It contains a vast collection of Python code snippets along with their corresponding instruction, input, output, and prompt information. By utilizing this dataset, users can gain valuable insights into various Python programming concepts and techniques.
The dataset is organized into columns to facilitate easy access to the required information. The instruction column holds the specific task or instruction that the Python code snippet is designed to perform. This allows users to understand the purpose or requirement of each code snippet at a glance.
The input column contains all necessary input data or parameters that are required for executing the Python code snippet accurately. These inputs provide context and enable users to comprehend how different variables or values impact the overall functioning of each code snippet.
Likewise, the output column presents expected results or outcomes that should be produced when executing each Python code snippet with its specified input values. This allows for validation and verification purposes, ensuring that each code snippet performs as intended.
In addition to instruction, input, and output details, this dataset also includes prompts. The prompt column provides additional context or information intended to assist users in better understanding the purpose or requirements of each particular Python code snippet.
By leveraging this comprehensive python_code_instructions_18k_alpaca training dataset, researchers and developers can delve into numerous real-world examples of Python programming challenges - helping them enhance their coding skills while gaining invaluable knowledge about effective implementation techniques across various domains
- Code Instruction Analysis: This dataset can be used to analyze different types of Python code instructions and identify patterns or common practices. Researchers or developers can use this dataset to gain insights into effective ways of writing code instructions.
- Code Output Prediction: With the given input and instruction, this dataset can be used to train models for predicting the expected output of a Python code snippet. This can be useful in automating the testing process or verifying the correctness of the code.
- Prompt Generation: Developers often struggle with providing clear and concise prompts for their code snippets. This dataset can serve as a resource for generating prompts by analyzing existing examples and extracting key information or requirements from them
If you use this dataset in your research, please credit the original authors. Data Source
License: CC0 1.0 Universal (CC0 1.0) - Public Domain Dedication No Copyright - You can copy, modify, distribute and perform the work, even for commercial purposes, all without asking permission. See Other Information.
File: train.csv | Column name | Description | |:----------------|:------------------------------------------------------------------------------------------------------------------| | instruction | Specific tasks or instructions assigned to each Python code snippet. (Text) | | input | The input data or parameters required for executing the code instruction. (Text) | | output | The expected result or output that should be produced when executing the code instruction. (Text) | | prompt | Additional information or context to help understand the purpose or requirements of each code instruction. (Text) |
If you use this dataset in your research, please credit the original authors. If you use this dataset in your research, please credit Tarun Bisht (From Huggingface).
Facebook
Twitterhttps://infobay.ai/terms-of-servicehttps://infobay.ai/terms-of-service
Algorithmic and repository-history code collections is an InfoBay corpus for enterprise AI teams that need traceable, expert-curated coding training data. A clearly separated corpus: a DSA collection for algorithmic training and a legacy-codebase collection for repository and software-engineering workflows.
Facebook
TwitterOpen Data Commons Attribution License (ODC-By) v1.0https://www.opendatacommons.org/licenses/by/1.0/
License information was derived automatically
This dataset was created during the Programming Language Ecosystem project from TU Wien using the code inside the repository https://github.com/ValentinFutterer/UsageOfProgramminglanguages2011-2023?tab=readme-ov-file.
The centerpiece of this repository is the usage_of_programming_languages_2011-2023.csv. This csv file shows the popularity of programming languages over the last 12 years in yearly increments. The repository also contains graphs created with the dataset. To get an accurate estimate on the popularity of programming languages, this dataset was created using 3 vastly different sources.
The dataset was created using the github repository above. As input data, three public datasets where used.
Taken from https://www.kaggle.com/datasets/pelmers/github-repository-metadata-with-5-stars/ by Peter Elmers. It is licensed under CC BY 4.0 https://creativecommons.org/licenses/by/4.0/. It shows metadata information (no code) of all github repositories with more than 5 stars.
Taken from https://github.com/pypl/pypl.github.io/tree/master, put online by the user pcarbonn. It is licensed under CC BY 3.0 https://creativecommons.org/licenses/by/3.0/. It shows from 2004 to 2023 for each month the share of programming related google searches per language.
Taken from https://insights.stackoverflow.com/survey. It is licensed under Open Data Commons Open Database License (ODbL) v1.0 https://opendatacommons.org/licenses/odbl/1-0/. It shows from 2011 to 2023 the results of the yearly stackoverflow developer survey.
All these datasets were downloaded on the 12.12.2023. The datasets are all in the github repository above
The dataset contains a column for the year and then many columns for the different languages, denoting their usage in percent. Additionally, vertical barcharts and piecharts for each year plus a line graph for each language over the whole timespan as png's are provided.
The languages that are going to be considered for the project can be seen here:
- Python
- C
- C++
- Java
- C#
- JavaScript
- PHP
- SQL
- Assembly
- Scratch
- Fortran
- Go
- Kotlin
- Delphi
- Swift
- Rust
- Ruby
- R
- COBOL
- F#
- Perl
- TypeScript
- Haskell
- Scala
This project is licensed under the Open Data Commons Open Database License (ODbL) v1.0 https://opendatacommons.org/licenses/odbl/1-0/ license.
TLDR: You are free to share, adapt, and create derivative works from this dataser as long as you attribute me, keep the database open (if you redistribute it), and continue to share-alike any adapted database under the ODbl.
Thanks go out to
- stackoverflow https://insights.stackoverflow.com/survey for providing the data from the yearly stackoverflow developer survey.
- the PYPL survey, https://github.com/pypl/pypl.github.io/tree/master for providing google search data.
- Peter Elmers, for crawling metadata on github repositories and providing the data https://www.kaggle.com/datasets/pelmers/github-repository-metadata-with-5-stars/.
Facebook
Twitterhttps://www.technavio.com/content/privacy-noticehttps://www.technavio.com/content/privacy-notice
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
Open source ranking dataset for Programming Language of China. Top Programming Language of China repositories on GitHub — ranked by stars, pull requests, issues & contributors. Compare trending projects and track growth over time.
Facebook
TwitterWeekly rankings of programming language demand by active job listings. Covers Python, SQL, TypeScript, Java, Go, Rust and more.
Facebook
Twitterhttps://www.wiseguyreports.com/pages/privacy-policyhttps://www.wiseguyreports.com/pages/privacy-policy
The Children's Programming Training Market was valued at USD 1165.4 Million in 2025 and is projected to grow to USD 3500 Million by 2035, at a CAGR of 11.7%. Children S Programming Training Market Overview: The Children's Programming Training Market Size was valued at 1,043.3 USD Million in 2024. The Children's Programming Training Market is expected to grow from 1,165.4 USD Million in 2025 to 3,500 USD Million by 2035. The Children's Programming Training Market CAGR (growth rate) is expected to be around 11.7% during the forecast period (2025 - 2035). Key Children S Programming Training Market Trends Highlighted The Global Children's Programming Training Market is witnessing significant shifts fueled by the increasing emphasis on early childhood education and the integration of technology in learning environments. As digital literacy becomes essential in modern curricula, there is a growing demand for programming training that aligns with children's cognitive development. This trend is encouraging educational institutions to incorporate coding and programming into their curricula, allowing children to develop critical problem-solving skills from a young age. Educational policies worldwide are increasingly focusing on enhancing digital competency, thus driving the adoption of programming training programs.Moreover, there exist substantial opportunities in the Global Children's Programming Training Market, particularly in underserved regions where access to quality educational resources is limited. Initiatives by governments and non-profit organizations aimed at bridging the digital divide are creating avenues for expansion and innovation. The advent of online learning platforms offers flexibility and accessibility, catering to diverse learning needs and ensuring that children from various backgrounds can participate in programming training. Recent trends also indicate a shift towards gamification and interactive learning as effective methods to engage younger audiences.This approach makes programming more appealing to children by transforming complex concepts into fun activities that promote engagement and retention. As parents and educators recognize the value of programming skills in future job markets, investment in children's programming training is expected to grow, shaping the future workforce. Thus, the Global Children's Programming Training Market is poised for significant growth and evolution as it adapts to increasingly dynamic educational landscapes. Source: Primary Research, Secondary Research, WGR Database and Analyst Review Children S Programming Training Market Segment Insights: Children S Programming Training Market Regional Insights The Global Children's Programming Training Market exhibits a diverse regional landscape with North America holding a majority stake, valued at 500 USD Million in 2024 and projected to reach 1,500 USD Million by 2035. This region is characterized by robust demand for children's programming education and training, driven by a strong emphasis on technology integration and engaging content. Europe is experiencing steady expansion as educational institutions adapt to the growing need for digital skills among children, fostering a dynamic environment for programming education.The APAC region is also witnessing moderate increase, reflecting the rising awareness about coding and programming skills in young learners, which aligns with initiatives by governments and private sectors alike. Meanwhile, South America is experiencing gradual growth, propelled by increasing access to digital platforms and educational resources. In the MEA region, challenges remain due to varying levels of educational infrastructure, but opportunities for growth are emerging as interest in technology education rises. Overall, the market growth across these regions is influenced by various factors including technological advancements, policy support, and changing educational paradigms that emphasize the importance of programming skills for the younger generation. Source: Primary Research, Secondary Research, WGR Database and Analyst Review North America : The Children's
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
NOTICE
This was done by me, someone with a learning Disability. So please do bare with me when updating this with more working data.
Multi-Language Programming Code Dataset
A curated dataset of original, non-scraped code examples across 7 programming environments: Python, JavaScript, Node.js, Java, C, C++, and Rust. The dataset ships in two parts that can be used separately or combined:
File Rows Description
code_dataset.jsonl / .csv 105 Hand-written… See the full description on the dataset page: https://huggingface.co/datasets/TGPRO32/Paragon-coding.
Facebook
TwitterSYNTHETIC-1
This is a subset of the task data used to construct SYNTHETIC-1. You can find the full collection here
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
This dataset tracks annual distribution of students across grade levels in Midview Virtual Programming
Facebook
Twitterhttps://www.mordorintelligence.com/terms-and-conditionshttps://www.mordorintelligence.com/terms-and-conditions
Complete dataset included in the full report. Detailed tables, regional splits, forecasts, and methodologies are available with purchase.
Facebook
Twitterhttps://www.verifiedmarketresearch.com/privacy-policy/https://www.verifiedmarketresearch.com/privacy-policy/
Programming Software Market size was valued at USD 38.44 Billion in 2025 and is projected to reach USD 90.60 Billion by 2033, growing at a CAGR of 11.61% from 2027 to 2033.The key market drivers for the Programming Software Market include the increasing adoption of cloud-native development, rising demand for AI-assisted coding tools, growing digital transformation initiatives across enterprises, expansion of DevOps and agile development practices, and the increasing need for secure, scalable, and efficient software development environments.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
Welcome to an exceptional dataset meticulously crafted for training state-of-the-art language models such as Gemma, Llama 2, Orca, and more.
Dataset Highlights - Challenging Questions : Immerse your language models in various Python programming questions designed to stimulate cognitive growth. - Real-world Inputs : Provide your models with authentic input scenarios, ensuring they are well-equipped to handle practical coding challenges. - Accurate Answers : Sharpen the precision of your language models by exposing them to meticulously crafted Python code solutions.
How to Get Started - Download : Grab a copy of the dataset and inject new life into your language models. - Build Brilliance : Watch your LLMs evolve as they engage with the challenging questions and nuanced coding scenarios. - Share & Collaborate : Join the Kaggle community to discuss, share insights, and collaborate with fellow enthusiasts.
Unleash the full potential of your language models with this dataset. Elevate your LLM training experience and witness unprecedented growth in language understanding and coding prowess. Happy coding !