Automated Organization Profile

Univ. Grenoble Alpes, CNRS, Grenoble INP, GIPSA-lab

Current S-Index

14.6

Sum of Dataset Indices for all datasets

Average Dataset Index per Dataset

1.5

Average Dataset Index per dataset

Total Datasets

10

Total datasets in this organization

Average FAIR Score

63.3%

Average FAIR Score per dataset

Total Citations

0

Total citations to the organization's datasets

Total Mentions

1

Total mentions of the organization's datasets

S-Index Interpretation

S-Index Over Time

Cumulative Citations Over Time

Cumulative Mentions Over Time

Datasets

PB2007 French acoustic-articulatory speech database

PB2007 acoustic-articulatory speech dataset Badin, P.,Bailly G., Ben Youssef A., Elisei F., Savariaux C., Hueber T.
Univ. Grenoble Alpes, CNRS, Grenoble INP, GIPSA-lab, 38000 Grenoble, France

LICENSE:
========
This dataset is made available under the Creative Commons Attribution Share-Alike (CC-BY-SA) license
CREDITS - ATTRIBUTION:
======================
If using this dataset, please cite one of the following studies (all of them exploit this dataset)
- Ben Youssef, A., Badin, P., Bailly, G. & Heracleous, P. (2009). Acoustic-to-articulatory inversion using speech recognition and trajectory formation based on phoneme hidden Markov models. In Interspeech 2009, vol., pp. 2255-2258. Brighton, UK.
- Ben Youssef, A., Badin, P. & Bailly, G. (2010). Can tongue be recovered from face? The answer of data-driven statistical models. In Interspeech 2010 (11th Annual Conference of the International Speech Communication Association) (T. Kobayashi, K. Hirose & S. Nakamura, editors), vol., pp. 2002-2005. Makuhari, Japan.
- Hueber T., Bailly G., Badin P., Elisei F., "Speaker Adaptation of an Acoustic-Articulatory Inversion Model
using Cascaded Gaussian Mixture Regressions", Proceedings of Interspeech, Lyon, France, 2013, pp. 2753-2757. DATA FILES DESCRIPTION:
=======================
/_seq/:
Electro-magnetic Articulography data, recorded at 100Hz
Sensors :
PAR01 : LT_x (lower incisor, x coordinate)
PAR02 : tip_x (tongue tip, x coordinate)
PAR03 : mid_x (tongue dorsum, x coordinate)
PAR04 : bck_x (tongue back, x coordinate)
PAR05 : LL_vis_x (lower lips, x coordinate)
PAR06 : UL_vis_x (upper lips, x coordinate)
PAR07 : LT_z (lower incisor, z coordinate)
PAR08 : tip_z (tongue tip, z coordinate)
PAR09 : mid_z (tongue dorsum, z coordinate)
PAR10 : bck_z (tongue back, z coordinate)
PAR11 : LL_vis_z (lower lips, z coordinate)
PAR12 : UL_vis_z (upper lips, z coordinate) /_wav16:
subject audio signal, synchronized with the EMA data
Format: PCA wav, 16kHz, 16bits /_lab: phonetic segmentation using the following set
__ (long pause), _ (short pause), a, e^ (as in "lait"), e (as in "blé"), i, y (as in "voiture"), u (as in "loup"), o^ (as in "pomme"),x (as in "pneu"), x^ (as in "coeur"), a~ (as in "flan"), e~ (as in "in"), x~ (as in "un"), o~ (as in "mon"), p, t, k, f, s, s^ (as in "CHat"), b, d, g, v, z, z^ (as in "les Gens"), m, n, r, l, w, h, j, o, q (schwa)

Authors

  • Badin, Pierre ;
  • Bailly, Gérard ;
  • Ben Youssef, Atef ;
  • Elisei, Frédéric ;
  • Savariaux, Christophe ;
  • Hueber, Thomas
0 Citations0 Mentions79% FAIR0.6 Dataset Index
10.5281/zenodo.63905972022

PB2007 French acoustic-articulatory speech database

PB2007 acoustic-articulatory speech dataset Badin, P.,Bailly G., Ben Youssef A., Elisei F., Savariaux C., Hueber T.
Univ. Grenoble Alpes, CNRS, Grenoble INP, GIPSA-lab, 38000 Grenoble, France

LICENSE:
========
This dataset is made available under the Creative Commons Attribution Share-Alike (CC-BY-SA) license
CREDITS - ATTRIBUTION:
======================
If using this dataset, please cite one of the following studies (all of them exploit this dataset)
- Ben Youssef, A., Badin, P., Bailly, G. & Heracleous, P. (2009). Acoustic-to-articulatory inversion using speech recognition and trajectory formation based on phoneme hidden Markov models. In Interspeech 2009, vol., pp. 2255-2258. Brighton, UK.
- Ben Youssef, A., Badin, P. & Bailly, G. (2010). Can tongue be recovered from face? The answer of data-driven statistical models. In Interspeech 2010 (11th Annual Conference of the International Speech Communication Association) (T. Kobayashi, K. Hirose & S. Nakamura, editors), vol., pp. 2002-2005. Makuhari, Japan.
- Hueber T., Bailly G., Badin P., Elisei F., "Speaker Adaptation of an Acoustic-Articulatory Inversion Model
using Cascaded Gaussian Mixture Regressions", Proceedings of Interspeech, Lyon, France, 2013, pp. 2753-2757. DATA FILES DESCRIPTION:
=======================
/_seq/:
Electro-magnetic Articulography data, recorded at 100Hz
Sensors :
PAR01 : LT_x (lower incisor, x coordinate)
PAR02 : tip_x (tongue tip, x coordinate)
PAR03 : mid_x (tongue dorsum, x coordinate)
PAR04 : bck_x (tongue back, x coordinate)
PAR05 : LL_vis_x (lower lips, x coordinate)
PAR06 : UL_vis_x (upper lips, x coordinate)
PAR07 : LT_z (lower incisor, z coordinate)
PAR08 : tip_z (tongue tip, z coordinate)
PAR09 : mid_z (tongue dorsum, z coordinate)
PAR10 : bck_z (tongue back, z coordinate)
PAR11 : LL_vis_z (lower lips, z coordinate)
PAR12 : UL_vis_z (upper lips, z coordinate) /_wav16:
subject audio signal, synchronized with the EMA data
Format: PCA wav, 16kHz, 16bits /_lab: phonetic segmentation using the following set
__ (long pause), _ (short pause), a, e^ (as in "lait"), e (as in "blé"), i, y (as in "voiture"), u (as in "loup"), o^ (as in "pomme"),x (as in "pneu"), x^ (as in "coeur"), a~ (as in "flan"), e~ (as in "in"), x~ (as in "un"), o~ (as in "mon"), p, t, k, f, s, s^ (as in "CHat"), b, d, g, v, z, z^ (as in "les Gens"), m, n, r, l, w, h, j, o, q (schwa)

Authors

  • Badin, Pierre ;
  • Bailly, Gérard ;
  • Ben Youssef, Atef ;
  • Elisei, Frédéric ;
  • Savariaux, Christophe ;
  • Hueber, Thomas
0 Citations0 Mentions73% FAIR0.6 Dataset Index
10.5281/zenodo.63905982022

PAirMax-Airbus (Version: 1.0)

This archive contains 5 panchromatic and multispectral bundles (at both both full and reduced resolution). These images are part of the PAirMax dataset (). This dataset is provided as a password protected folder as data can be accessed only after accepting the Airbus license. Description: The images were derived from original acquisitions by the Pléiades and Spot7 satellites and are provided courtesy of Airbus. The original images full scenes can be accessed at: https://sandbox.intelligence-airbusds.com -> Pansharpening dataset. The 5 images in this archive are listed below. Please refer to [1] for more details on the images and the preprocessing done. - Pl_Hous_Urb - Pl_Sacr_Mix - Pl_Stoc_Urb - S7_Napl_Urb - S7_NewY_Mix Instruction for retrieving the password: - Go to https://sandbox.intelligence-airbusds.com - Fill the form for requesting the Pansharpening dataset (need to accept the Airbus license) - The password will be provided in the confirmation email. ---------------------------------------------------------------------------- () The PAirMax dataset is a collection of data with the aim of assessing the performance of pansharpening algorithms. The data collection includes 5 test cases selected at full resolution (FR), acquired by two sensors belonging to the Airbus' constellation of high-resolution imaging satellites. Moreover, 5 related test cases at reduced resolution (RR), simulated according to the Wald’s protocol, are included, thus resulting in 10 challenging test cases for pansharpening performance assessment. For further details, please, refer to the paper: [1] G. Vivone, M. Dalla Mura, A. Garzelli, and F. Pacifici, "A Benchmarking Protocol for Pansharpening: Dataset, Pre-processing, and Quality Assessment," IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2021.

Authors

  • Dalla Mura, Mauro
0 Citations0 Mentions15% FAIR0.1 Dataset Index
10.18709/perscido.2021.09.ds3532021

TMD-CAPTIMOVE (Version: 1.0)

This database called "TMD-CAPTIMOVE" provides transportation mode labelled data collected by 34 volunteers for a total time duration of around 48 hours. Considered transportation modes are: On-foot (Walking, Stairs, Elevators), Bike, Scooter, Bus, Tram. The number of labels is 11: Still, Walk, Upstairs, Downstairs, Elevator up, Elevator down, Bike, Electric scooter, kick scooter, Bus, Tram. Sensor data are: Acceleration (m/s²), angular rate (°/s), atmospheric pressure (hPa), heart rate (beat per minute (BPM)). The sampling frequency for all data is 32 Hertz.

Authors

  • Fourati, Hassen
0 Citations0 Mentions15% FAIR0.1 Dataset Index
10.18709/perscido.2020.04.ds3102020

Deeppredspeech: Computational Models Of Predictive Speech Coding Based On Deep Learning

This dataset contains all data, source code, pre-trained computational predictive models and experimental results related to: Hueber T., Tatulli E., Girin L., Schwatz, J-L "How predictive can be predictions in the neurocognitive processing of auditory and audiovisual speech? A deep learning study." (biorXiv preprint https://doi.org/10.1101/471581). Raw data are extracted from the publicly available database NTCD-TIMIT (10.5281/zenodo.260228). Audio recordings are available in the audio_clean/ directoryPost-processed lip image sequences are available in the lips_roi/ directory (67x67 pixels, 8bits, obtained by lossless inverse DCT-2D transform from the DCT feature available in the original repository of NTCD-TIMIT)Phonetic segmentation (extracted from NTCD-TIMIT original zenodo repository) is available in the HTK MLF file volunteer_labelfiles.mlfAudio features (MFCC-spectrogram and log-spectrogram) are available in the mfcc_16k/ and fft_16k/ directories. Models (audio-only, video-only and audiovisual, based on deep feed-forward neural networks and/or convolutional neural network, in .h5 format, trained with Keras 2.0 toolkit) and data normalization parameters (in .dat scikit-learn format) are available in models_mfcc/ and models_logspectro/ directoriesPredicted and target (ground truth) MFCC-spectro (resp. log-spectro) for the test databases (1909 sentences), and for the different values of (\tau_p) or (\tau_f) are available in pred_testdb_mfccspectro/ (resp. pred_testdb_logspectro/) directorySource code for extracting audio features, training and evaluating the models is available on GitHub https://github.com/thueber/DeepPredSpeech/All directories have been zipped before upload.Feel free to contact me for more details.Thomas Hueber, Ph. D., CNRS research fellow, GIPSA-lab, Grenoble, France, [email protected]

Authors

  • Hueber, Thomas ;
  • Tatulli, Eric ;
  • Girin, Laurent ;
  • Schwartz, Jean-Luc
0 Citations1 Mention73% FAIR1.0 Dataset Index
10.5281/zenodo.14879742018

DeepPredSpeech: computational models of predictive speech coding based on deep learning (Version: 1.0)

This dataset contains all data, source code, pre-trained computational predictive models and experimental results related to: Hueber T., Tatulli E., Girin L., Schwatz, J-L "How predictive can be predictions in the neurocognitive processing of auditory and audiovisual speech? A deep learning study." (biorXiv preprint https://doi.org/10.1101/471581). Raw data are extracted from the publicly available database NTCD-TIMIT (10.5281/zenodo.260228). Audio recordings are available in the audio_clean/ directory Post-processed lip image sequences are available in the lips_roi/ directory (67x67 pixels, 8bits, obtained by lossless inverse DCT-2D transform from the DCT feature available in the original repository of NTCD-TIMIT) Phonetic segmentation (extracted from NTCD-TIMIT original zenodo repository) is available in the HTK MLF file volunteer_labelfiles.mlf Audio features (MFCC-spectrogram and log-spectrogram) are available in the mfcc_16k/ and fft_16k/ directories. Models (audio-only, video-only and audiovisual, based on deep feed-forward neural networks and/or convolutional neural network, in .h5 format, trained with Keras 2.0 toolkit) and data normalization parameters (in .dat scikit-learn format) are available in models_mfcc/ and models_logspectro/ directories Predicted and target (ground truth) MFCC-spectro (resp. log-spectro) for the test databases (1909 sentences), and for the different values of (\tau_p) or (\tau_f) are available in pred_testdb_mfccspectro/ (resp. pred_testdb_logspectro/) directory Source code for extracting audio features, training and evaluating the models is available on GitHub https://github.com/thueber/DeepPredSpeech/ All directories have been zipped before upload. Feel free to contact me for more details. Thomas Hueber, Ph. D., CNRS research fellow, GIPSA-lab, Grenoble, France, [email protected]

Authors

  • Hueber, Thomas ;
  • Tatulli, Eric ;
  • Girin, Laurent ;
  • Schwartz, Jean-Luc
0 Citations0 Mentions73% FAIR0.4 Dataset Index
10.5281/zenodo.14879732018

DeepPredSpeech: computational models of predictive speech coding based on deep learning (Version: 1.0)

This dataset contains all data, source code, pre-trained computational predictive models and experimental results related to: Hueber T., Tatulli E., Girin L., Schwatz, J-L "How predictive can be predictions in the neurocognitive processing of auditory and audiovisual speech? A deep learning study." (biorXiv preprint https://doi.org/10.1101/471581). Raw data are extracted from the publicly available database NTCD-TIMIT (10.5281/zenodo.260228). Audio recordings are available in the audio_clean/ directory Post-processed lip image sequences are available in the lips_roi/ directory (67x67 pixels, 8bits, obtained by lossless inverse DCT-2D transform from the DCT feature available in the original repository of NTCD-TIMIT) Phonetic segmentation (extracted from NTCD-TIMIT original zenodo repository) is available in the HTK MLF file volunteer_labelfiles.mlf Audio features (MFCC-spectrogram and log-spectrogram) are available in the mfcc_16k/ and fft_16k/ directories. Models (audio-only, video-only and audiovisual, based on deep feed-forward neural networks and/or convolutional neural network, in .h5 format, trained with Keras 2.0 toolkit) and data normalization parameters (in .dat scikit-learn format) are available in models_mfcc/ and models_logspectro/ directories Predicted and target (ground truth) MFCC-spectro (resp. log-spectro) for the test databases (1909 sentences), and for the different values of (\tau_p) or (\tau_f) are available in pred_testdb_mfccspectro/ (resp. pred_testdb_logspectro/) directory Source code for extracting audio features, training and evaluating the models is available on GitHub https://github.com/thueber/DeepPredSpeech/ All directories have been zipped before upload. Feel free to contact me for more details. Thomas Hueber, Ph. D., CNRS research fellow, GIPSA-lab, Grenoble, France, [email protected]

Authors

  • Hueber, Thomas ;
  • Tatulli, Eric ;
  • Girin, Laurent ;
  • Schwartz, Jean-Luc
0 Citations0 Mentions73% FAIR0.4 Dataset Index
10.5281/zenodo.35280682018

CSF18 (Version: 1.1)

CSF18 - Multimodal database of French Cued-speech (revised in 2022) Dataset used in "Visual recognition of continuous Cued Speech using a tandem CNN-HMM approach", by Liu, Hueber, Feng, Beautemps (submitted to Interspeech 2018) 476 sentences (i.e. 2 repetitions of 238 sentences) uttered by a professional French Cued-speech coder video/ PNG images, 576x720, 50fps (after deinterleave) audio/ WAV, 16kHz, 16bits prompt.txt: Text prompt of the recorded sentences corpus_mlf.txt: Phonetic transcription aligned on the audio signal (HTK format, Master Label File) obtained using the LiaPhon phonetizer and a forced-alignment HMM-based procedure (no manual check) corpus_mlf_updated_icassp2022.txt: Manually checked/cleaned version of corpus_mlf.txt (see Sankar et al., ICASSP 2022 paper) phonelist.txt: list of the 34 labels used to encode French phonemes at GIPSA-lab.

Authors

  • Liu, Li ;
  • Hueber, Thomas ;
  • Feng, Gang ;
  • Beautemps, Denis ;
  • Sankar, Sanjana
0 Citations0 Mentions77% FAIR0.4 Dataset Index
10.5281/zenodo.55548492018

CSF18 (Version: 1.1)

CSF18 - Multimodal database of French Cued-speech (revised in 2022) Dataset used in "Visual recognition of continuous Cued Speech using a tandem CNN-HMM approach", by Liu, Hueber, Feng, Beautemps (submitted to Interspeech 2018) 476 sentences (i.e. 2 repetitions of 238 sentences) uttered by a professional French Cued-speech coder video/ PNG images, 576x720, 50fps (after deinterleave) audio/ WAV, 16kHz, 16bits prompt.txt: Text prompt of the recorded sentences corpus_mlf.txt: Phonetic transcription aligned on the audio signal (HTK format, Master Label File) obtained using the LiaPhon phonetizer and a forced-alignment HMM-based procedure (no manual check) corpus_mlf_updated_icassp2022.txt: Manually checked/cleaned version of corpus_mlf.txt (see Sankar et al., ICASSP 2022 paper) phonelist.txt: list of the 34 labels used to encode French phonemes at GIPSA-lab.

Authors

  • Liu, Li ;
  • Hueber, Thomas ;
  • Feng, Gang ;
  • Beautemps, Denis ;
  • Sankar, Sanjana
0 Citations0 Mentions77% FAIR0.4 Dataset Index
10.5281/zenodo.12060002018

Csf18

CSF18 - Multimodal database of French Cued-speech Dataset used in "Visual recognition of continuous Cued Speech using a tandem CNN-HMM approach", by Liu, Hueber, Feng, Beautemps (submitted to Interspeech 2018)476 sentences (i.e. 2 repetitions of 238 sentences) uttered by a professional French Cued-speech codervideo/ PNG images, 576x720, 50fps (after deinterleave)audio/ WAV, 16kHz, 16bitsprompt.txt: Text prompt of the recorded sentences corpus_mlf.txt: Phonetic transcription aligned on the audio signal (HTK format, Master Label File) obtained using the LiaPhon phonetizer and a forced-alignment HMM-based procedure (no manual check)phonelist.txt: list of the 34 labels used to encode French phonemes at GIPSA-lab.

Authors

  • Liu, Li ;
  • Hueber, Thomas ;
  • Feng, Gang ;
  • Beautemps, Denis
0 Citations0 Mentions77% FAIR0.4 Dataset Index
10.5281/zenodo.12060012018