3 datasets found

F
Filipino General Conversation Speech Dataset for ASR
futurebeeai.com
wav
Updated Aug 1, 2022
+ more versions
Share
Facebook
Twitter
Email
Click to copy link
Link copied
Cite
FutureBee AI (2022). Filipino General Conversation Speech Dataset for ASR [Dataset]. https://www.futurebeeai.com/dataset/speech-dataset/general-conversation-filipino-philippines
Explore at:
wavAvailable download formats
Dataset updated
Aug 1, 2022
Dataset provided by
FutureBeeAI
Authors
FutureBee AI
License
https://www.futurebeeai.com/policies/ai-data-license-agreementhttps://www.futurebeeai.com/policies/ai-data-license-agreement
Area covered
Philippines
Dataset funded by
FutureBeeAI
Description
Introduction
Welcome to the Filipino General Conversation Speech Dataset — a rich, linguistically diverse corpus purpose-built to accelerate the development of Filipino speech technologies. This dataset is designed to train and fine-tune ASR systems, spoken language understanding models, and generative voice AI tailored to real-world Filipino communication.
Curated by FutureBeeAI, this 30 hours dataset offers unscripted, spontaneous two-speaker conversations across a wide array of real-life topics. It enables researchers, AI developers, and voice-first product teams to build robust, production-grade Filipino speech models that understand and respond to authentic Filipino accents and dialects.
Speech Data
The dataset comprises 30 hours of high-quality audio, featuring natural, free-flowing dialogue between native speakers of Filipino. These sessions range from informal daily talks to deeper, topic-specific discussions, ensuring variability and context richness for diverse use cases.
•Participant Diversity:
•
Speakers: 60 verified native Filipino speakers from FutureBeeAI’s contributor community.

•
Regions: Representing various provinces of Philippines to ensure dialectal diversity and demographic balance.

•
Demographics: A balanced gender ratio (60% male, 40% female) with participant ages ranging from 18 to 70 years.

•Recording Details:
•
Conversation Style: Unscripted, spontaneous peer-to-peer dialogues.

•
Duration: Each conversation ranges from 15 to 60 minutes.

•
Audio Format: Stereo WAV files, 16-bit depth, recorded at 16kHz sample rate.

•
Environment: Quiet, echo-free settings with no background noise.

Topic Diversity
The dataset spans a wide variety of everyday and domain-relevant themes. This topic diversity ensures the resulting models are adaptable to broad speech contexts.
•Sample Topics Include:
•Family & Relationships
•Food & Recipes
•Education & Career
•Healthcare Discussions
•Social Issues
•Technology & Gadgets
•Travel & Local Culture
•Shopping & Marketplace Experiences, and many more.
Transcription
Each audio file is paired with a human-verified, verbatim transcription available in JSON format.
•Transcription Highlights:
•Speaker-segmented dialogues
•Time-coded utterances
•Non-speech elements (pauses, laughter, etc.)
•High transcription accuracy, achieved through double QA pass, average WER < 5%
These transcriptions are production-ready, enabling seamless integration into ASR model pipelines or conversational AI workflows.
Metadata
The dataset comes with granular metadata for both speakers and recordings:
•
Speaker Metadata: Age, gender, accent, dialect, state/province, and participant ID.

•
Recording Metadata: Topic, duration, audio format, device type, and sample rate.

Such metadata helps developers fine-tune model training and supports use-case-specific filtering or demographic analysis.
Usage and Applications
This dataset is a versatile resource for multiple Filipino speech and language AI applications:
•
ASR Development: Train accurate speech-to-text systems for Filipino.

•
Voice Assistants: Build smart assistants capable of understanding natural Filipino conversations.
F
Filipino TTS Speech Dataset for Speech Synthesis
futurebeeai.com
wav
Updated Aug 1, 2022
+ more versions
Share
Facebook
Twitter
Email
Click to copy link
Link copied
Cite
FutureBee AI (2022). Filipino TTS Speech Dataset for Speech Synthesis [Dataset]. https://www.futurebeeai.com/dataset/speech-dataset/tts-monolgue-filipino-philippines
Explore at:
wavAvailable download formats
Dataset updated
Aug 1, 2022
Dataset provided by
FutureBeeAI
Authors
FutureBee AI
License
https://www.futurebeeai.com/policies/ai-data-license-agreementhttps://www.futurebeeai.com/policies/ai-data-license-agreement
Dataset funded by
FutureBeeAI
Description
The Filipino TTS Monologue Speech Dataset is a professionally curated resource built to train realistic, expressive, and production-grade text-to-speech (TTS) systems. It contains studio-recorded long-form speech by trained native Filipino voice artists, each contributing 1 to 2 hours of clean, uninterrupted monologue audio.
Unlike typical prompt-based datasets with short, isolated phrases, this collection features long-form, topic-driven monologues that mirror natural human narration. It includes content types that are directly useful for real-world applications, like audiobook-style storytelling, educational lectures, health advisories, product explainers, digital how-tos, formal announcements, and more.
All recordings are captured in professional studios using high-end equipment and under the guidance of experienced voice directors.
Recording & Audio Quality
•
Audio Format: WAV, 48 kHz, available in 16-bit, 24-bit, and 32-bit depth

•
SNR: Minimum 30 dB

•
Channel: Mono

•
Recording Duration: 20-30 minutes

•
Recording Environment: Studio-controlled, acoustically treated rooms

•
Per Speaker Volume: 1–2 hours of speech per artist

•
Quality Control: Each file is reviewed and cleaned for common acoustic issues, including: reverberation, lip smacks, mouth clicks, thumping, hissing, plosives, sibilance, background noise, static interference, clipping, and other artifacts.

Only clean, production-grade audio makes it into the final dataset.
Voice Artist Selection
All voice artists are native Filipino speakers with professional training or prior experience in narration. We ensure a diverse pool in terms of age, gender, and region to bring a balanced and rich vocal dataset.
•Artist Profile:
•Gender: Male and Female
•Age Range: 20–60 years
•Regions: Native Filipino-speaking states from Philippines
•
Selection Process: All artists are screened, onboarded, and sample-approved using FutureBeeAI’s proprietary Yugo platform.

Script Quality & Coverage
Scripts are not generic or repetitive. Scripts are professionally authored by domain experts to reflect real-world use cases. They avoid redundancy and include modern vocabulary, emotional range, and phonetically rich sentence structures.
•
Word Count per Script: 3,000–5,000 words per 30-minute session

•Content Types:
•Storytelling
•Script and book reading
•Informational explainers
•Government service instructions
•E-commerce tutorials
•Motivational content
•Health & wellness guides
•Education & career advice
•
Linguistic Design: Balanced punctuation, emotional range, modern syntax, and vocabulary diversity

Transcripts & Alignment
While the script is used during the recording, we also provide post-recording updates to ensure the transcript reflects the final spoken audio. Minor edits are made to adjust for skipped or rephrased words.
•
Segmentation: Time-stamped at the sentence level, aligned to actual spoken delivery

•
Format: Available in plain text and JSON

•Post-processing:
•Corrected for
Data from: Lost on the frontline, and lost in the data: COVID-19 deaths...
figshare.com
zip
Updated Jul 22, 2022
Share
Facebook
Twitter
Email
Click to copy link
Link copied
Cite
Loraine Escobedo (2022). Lost on the frontline, and lost in the data: COVID-19 deaths among Filipinx healthcare workers in the United States [Dataset]. http://doi.org/10.6084/m9.figshare.20353368.v1
Explore at:
zipAvailable download formats
Unique identifier
https://doi.org/10.6084/m9.figshare.20353368.v1
Dataset updated
Jul 22, 2022
Dataset provided by
Figsharehttp://figshare.com/
figshare
Authors
Loraine Escobedo
License
Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
Area covered
United States
Description
To estimate county of residence of Filipinx healthcare workers who died of COVID-19, we retrieved data from the Kanlungan website during the month of December 2020.22 In deciding who to include on the website, the AF3IRM team that established the Kanlungan website set two standards in data collection. First, the team found at least one source explicitly stating that the fallen healthcare worker was of Philippine ancestry; this was mostly media articles or obituaries sharing the life stories of the deceased. In a few cases, the confirmation came directly from the deceased healthcare worker's family member who submitted a tribute. Second, the team required a minimum of two sources to identify and announce fallen healthcare workers. We retrieved 86 US tributes from Kanlungan, but only 81 of them had information on county of residence. In total, 45 US counties with at least one reported tribute to a Filipinx healthcare worker who died of COVID-19 were identified for analysis and will hereafter be referred to as “Kanlungan counties.” Mortality data by county, race, and ethnicity came from the National Center for Health Statistics (NCHS).24 Updated weekly, this dataset is based on vital statistics data for use in conducting public health surveillance in near real time to provide provisional mortality estimates based on data received and processed by a specified cutoff date, before data are finalized and publicly released.25 We used the data released on December 30, 2020, which included provisional COVID-19 death counts from February 1, 2020 to December 26, 2020—during the height of the pandemic and prior to COVID-19 vaccines being available—for counties with at least 100 total COVID-19 deaths. During this time period, 501 counties (15.9% of the total 3,142 counties in all 50 states and Washington DC)26 met this criterion. Data on COVID-19 deaths were available for six major racial/ethnic groups: Non-Hispanic White, Non-Hispanic Black, Non-Hispanic Native Hawaiian or Other Pacific Islander, Non-Hispanic American Indian or Alaska Native, Non-Hispanic Asian (hereafter referred to as Asian American), and Hispanic. People with more than one race, and those with unknown race were included in the “Other” category. NCHS suppressed county-level data by race and ethnicity if death counts are less than 10. In total, 133 US counties reported COVID-19 mortality data for Asian Americans. These data were used to calculate the percentage of all COVID-19 decedents in the county who were Asian American. We used data from the 2018 American Community Survey (ACS) five-year estimates, downloaded from the Integrated Public Use Microdata Series (IPUMS) to create county-level population demographic variables.27 IPUMS is publicly available, and the database integrates samples using ACS data from 2000 to the present using a high degree of precision.27 We applied survey weights to calculate the following variables at the county-level: median age among Asian Americans, average income to poverty ratio among Asian Americans, the percentage of the county population that is Filipinx, and the percentage of healthcare workers in the county who are Filipinx. Healthcare workers encompassed all healthcare practitioners, technical occupations, and healthcare service occupations, including nurse practitioners, physicians, surgeons, dentists, physical therapists, home health aides, personal care aides, and other medical technicians and healthcare support workers. County-level data were available for 107 out of the 133 counties (80.5%) that had NCHS data on the distribution of COVID-19 deaths among Asian Americans, and 96 counties (72.2%) with Asian American healthcare workforce data. The ACS 2018 five-year estimates were also the source of county-level percentage of the Asian American population (alone or in combination) who are Filipinx.8 In addition, the ACS provided county-level population counts26 to calculate population density (people per 1,000 people per square mile), estimated by dividing the total population by the county area, then dividing by 1,000 people. The county area was calculated in ArcGIS 10.7.1 using the county boundary shapefile and projected to Albers equal area conic (for counties in the US contiguous states), Hawai’i Albers Equal Area Conic (for Hawai’i counties), and Alaska Albers Equal Area Conic (for Alaska counties).20
Not seeing a result you expected?
Learn how you can add new datasets to our index.

Facebook

Twitter

Click to copy link

Link copied

Cite

FutureBee AI (2022). Filipino General Conversation Speech Dataset for ASR [Dataset]. https://www.futurebeeai.com/dataset/speech-dataset/general-conversation-filipino-philippines

Filipino General Conversation Speech Dataset for ASR

Filipino General Conversation Speech Corpus

Explore at:

wavAvailable download formats

Dataset updated

Aug 1, 2022

Dataset provided by

FutureBeeAI

Authors

FutureBee AI

License

https://www.futurebeeai.com/policies/ai-data-license-agreementhttps://www.futurebeeai.com/policies/ai-data-license-agreement

Area covered

Philippines

Dataset funded by

FutureBeeAI

Description

Introduction

Welcome to the Filipino General Conversation Speech Dataset — a rich, linguistically diverse corpus purpose-built to accelerate the development of Filipino speech technologies. This dataset is designed to train and fine-tune ASR systems, spoken language understanding models, and generative voice AI tailored to real-world Filipino communication.

Curated by FutureBeeAI, this 30 hours dataset offers unscripted, spontaneous two-speaker conversations across a wide array of real-life topics. It enables researchers, AI developers, and voice-first product teams to build robust, production-grade Filipino speech models that understand and respond to authentic Filipino accents and dialects.

Speech Data

The dataset comprises 30 hours of high-quality audio, featuring natural, free-flowing dialogue between native speakers of Filipino. These sessions range from informal daily talks to deeper, topic-specific discussions, ensuring variability and context richness for diverse use cases.

•Participant Diversity:

•

Speakers: 60 verified native Filipino speakers from FutureBeeAI’s contributor community.

•

Regions: Representing various provinces of Philippines to ensure dialectal diversity and demographic balance.

•

Demographics: A balanced gender ratio (60% male, 40% female) with participant ages ranging from 18 to 70 years.

•Recording Details:

•

Conversation Style: Unscripted, spontaneous peer-to-peer dialogues.

•

Duration: Each conversation ranges from 15 to 60 minutes.

•

Audio Format: Stereo WAV files, 16-bit depth, recorded at 16kHz sample rate.

•

Environment: Quiet, echo-free settings with no background noise.

Topic Diversity

The dataset spans a wide variety of everyday and domain-relevant themes. This topic diversity ensures the resulting models are adaptable to broad speech contexts.

•Sample Topics Include:

•Family & Relationships

•Food & Recipes

•Education & Career

•Healthcare Discussions

•Social Issues

•Technology & Gadgets

•Travel & Local Culture

•Shopping & Marketplace Experiences, and many more.

Transcription

Each audio file is paired with a human-verified, verbatim transcription available in JSON format.

•Transcription Highlights:

•Speaker-segmented dialogues

•Time-coded utterances

•Non-speech elements (pauses, laughter, etc.)

•High transcription accuracy, achieved through double QA pass, average WER < 5%

These transcriptions are production-ready, enabling seamless integration into ASR model pipelines or conversational AI workflows.

Metadata

The dataset comes with granular metadata for both speakers and recordings:

•

Speaker Metadata: Age, gender, accent, dialect, state/province, and participant ID.

•

Recording Metadata: Topic, duration, audio format, device type, and sample rate.

Such metadata helps developers fine-tune model training and supports use-case-specific filtering or demographic analysis.

Usage and Applications

This dataset is a versatile resource for multiple Filipino speech and language AI applications:

•

ASR Development: Train accurate speech-to-text systems for Filipino.

•

Voice Assistants: Build smart assistants capable of understanding natural Filipino conversations.

Clear search

Close search

Google apps

Main menu

Filipino General Conversation Speech Dataset for ASR

Introduction

Speech Data

Topic Diversity

Transcription

Metadata

Usage and Applications

Filipino TTS Speech Dataset for Speech Synthesis

Recording & Audio Quality

Voice Artist Selection

Script Quality & Coverage

Transcripts & Alignment

Data from: Lost on the frontline, and lost in the data: COVID-19 deaths...

Filipino General Conversation Speech Dataset for ASRSee More Versions

Filipino General Conversation Speech Corpus

Introduction

Speech Data

Topic Diversity

Transcription

Metadata

Usage and Applications

Filipino General Conversation Speech Dataset for ASR