Automated Organization ProfileUniv. Grenoble Alpes, CNRS, Grenoble INP, GIPSA-lab
Univ. Grenoble Alpes, CNRS, Grenoble INP, GIPSA-lab
Current S-Index
Sum of Dataset Indices for all datasets
Average Dataset Index per Dataset
Average Dataset Index per dataset
Total Datasets
Total datasets in this organization
Average FAIR Score
Average FAIR Score per dataset
Total Citations
Total citations to the organization's datasets
Total Mentions
Total mentions of the organization's datasets
S-Index Interpretation
The S-Index (Sharing Index) is a comprehensive metric that represents the cumulative impact of all your datasets. It is calculated as the sum of Dataset Index scores across all your claimed datasets.
What it means:
- A higher S-index indicates greater overall impact of your datasets relative to typical datasets in their fields of research
- The S-Index grows as you add more datasets or as existing datasets gain more citations and mentions
- It provides a single number to track your research data impact over time
Current S-Index: 14.6 (sum of 10 datasets Dataset Index scores)
More information here.
S-Index Over Time
Cumulative Citations Over Time
Cumulative Mentions Over Time
Datasets
PB2007 acoustic-articulatory speech dataset Badin, P.,Bailly G., Ben Youssef A., Elisei F., Savariaux C., Hueber T.
Univ. Grenoble Alpes, CNRS, Grenoble INP, GIPSA-lab, 38000 Grenoble, France
LICENSE:
========
This dataset is made available under the Creative Commons Attribution Share-Alike (CC-BY-SA) license
CREDITS - ATTRIBUTION:
======================
If using this dataset, please cite one of the following studies (all of them exploit this dataset)
- Ben Youssef, A., Badin, P., Bailly, G. & Heracleous, P. (2009). Acoustic-to-articulatory inversion using speech recognition and trajectory formation based on phoneme hidden Markov models. In Interspeech 2009, vol., pp. 2255-2258. Brighton, UK.
- Ben Youssef, A., Badin, P. & Bailly, G. (2010). Can tongue be recovered from face? The answer of data-driven statistical models. In Interspeech 2010 (11th Annual Conference of the International Speech Communication Association) (T. Kobayashi, K. Hirose & S. Nakamura, editors), vol., pp. 2002-2005. Makuhari, Japan.
- Hueber T., Bailly G., Badin P., Elisei F., "Speaker Adaptation of an Acoustic-Articulatory Inversion Model
using Cascaded Gaussian Mixture Regressions", Proceedings of Interspeech, Lyon, France, 2013, pp. 2753-2757. DATA FILES DESCRIPTION:
=======================
/_seq/:
Electro-magnetic Articulography data, recorded at 100Hz
Sensors :
PAR01 : LT_x (lower incisor, x coordinate)
PAR02 : tip_x (tongue tip, x coordinate)
PAR03 : mid_x (tongue dorsum, x coordinate)
PAR04 : bck_x (tongue back, x coordinate)
PAR05 : LL_vis_x (lower lips, x coordinate)
PAR06 : UL_vis_x (upper lips, x coordinate)
PAR07 : LT_z (lower incisor, z coordinate)
PAR08 : tip_z (tongue tip, z coordinate)
PAR09 : mid_z (tongue dorsum, z coordinate)
PAR10 : bck_z (tongue back, z coordinate)
PAR11 : LL_vis_z (lower lips, z coordinate)
PAR12 : UL_vis_z (upper lips, z coordinate) /_wav16:
subject audio signal, synchronized with the EMA data
Format: PCA wav, 16kHz, 16bits /_lab: phonetic segmentation using the following set
__ (long pause), _ (short pause), a, e^ (as in "lait"), e (as in "blé"), i, y (as in "voiture"), u (as in "loup"), o^ (as in "pomme"),x (as in "pneu"), x^ (as in "coeur"), a~ (as in "flan"), e~ (as in "in"), x~ (as in "un"), o~ (as in "mon"), p, t, k, f, s, s^ (as in "CHat"), b, d, g, v, z, z^ (as in "les Gens"), m, n, r, l, w, h, j, o, q (schwa)
Authors
- Badin, Pierre ;
- Bailly, Gérard ;
- Ben Youssef, Atef ;
- Elisei, Frédéric ;
- Savariaux, Christophe ;
- Hueber, Thomas
PB2007 acoustic-articulatory speech dataset Badin, P.,Bailly G., Ben Youssef A., Elisei F., Savariaux C., Hueber T.
Univ. Grenoble Alpes, CNRS, Grenoble INP, GIPSA-lab, 38000 Grenoble, France
LICENSE:
========
This dataset is made available under the Creative Commons Attribution Share-Alike (CC-BY-SA) license
CREDITS - ATTRIBUTION:
======================
If using this dataset, please cite one of the following studies (all of them exploit this dataset)
- Ben Youssef, A., Badin, P., Bailly, G. & Heracleous, P. (2009). Acoustic-to-articulatory inversion using speech recognition and trajectory formation based on phoneme hidden Markov models. In Interspeech 2009, vol., pp. 2255-2258. Brighton, UK.
- Ben Youssef, A., Badin, P. & Bailly, G. (2010). Can tongue be recovered from face? The answer of data-driven statistical models. In Interspeech 2010 (11th Annual Conference of the International Speech Communication Association) (T. Kobayashi, K. Hirose & S. Nakamura, editors), vol., pp. 2002-2005. Makuhari, Japan.
- Hueber T., Bailly G., Badin P., Elisei F., "Speaker Adaptation of an Acoustic-Articulatory Inversion Model
using Cascaded Gaussian Mixture Regressions", Proceedings of Interspeech, Lyon, France, 2013, pp. 2753-2757. DATA FILES DESCRIPTION:
=======================
/_seq/:
Electro-magnetic Articulography data, recorded at 100Hz
Sensors :
PAR01 : LT_x (lower incisor, x coordinate)
PAR02 : tip_x (tongue tip, x coordinate)
PAR03 : mid_x (tongue dorsum, x coordinate)
PAR04 : bck_x (tongue back, x coordinate)
PAR05 : LL_vis_x (lower lips, x coordinate)
PAR06 : UL_vis_x (upper lips, x coordinate)
PAR07 : LT_z (lower incisor, z coordinate)
PAR08 : tip_z (tongue tip, z coordinate)
PAR09 : mid_z (tongue dorsum, z coordinate)
PAR10 : bck_z (tongue back, z coordinate)
PAR11 : LL_vis_z (lower lips, z coordinate)
PAR12 : UL_vis_z (upper lips, z coordinate) /_wav16:
subject audio signal, synchronized with the EMA data
Format: PCA wav, 16kHz, 16bits /_lab: phonetic segmentation using the following set
__ (long pause), _ (short pause), a, e^ (as in "lait"), e (as in "blé"), i, y (as in "voiture"), u (as in "loup"), o^ (as in "pomme"),x (as in "pneu"), x^ (as in "coeur"), a~ (as in "flan"), e~ (as in "in"), x~ (as in "un"), o~ (as in "mon"), p, t, k, f, s, s^ (as in "CHat"), b, d, g, v, z, z^ (as in "les Gens"), m, n, r, l, w, h, j, o, q (schwa)
Authors
- Badin, Pierre ;
- Bailly, Gérard ;
- Ben Youssef, Atef ;
- Elisei, Frédéric ;
- Savariaux, Christophe ;
- Hueber, Thomas
This archive contains 5 panchromatic and multispectral bundles (at both both full and reduced resolution). These images are part of the PAirMax dataset (). This dataset is provided as a password protected folder as data can be accessed only after accepting the Airbus license. Description: The images were derived from original acquisitions by the Pléiades and Spot7 satellites and are provided courtesy of Airbus. The original images full scenes can be accessed at: https://sandbox.intelligence-airbusds.com -> Pansharpening dataset. The 5 images in this archive are listed below. Please refer to [1] for more details on the images and the preprocessing done. - Pl_Hous_Urb - Pl_Sacr_Mix - Pl_Stoc_Urb - S7_Napl_Urb - S7_NewY_Mix Instruction for retrieving the password: - Go to https://sandbox.intelligence-airbusds.com - Fill the form for requesting the Pansharpening dataset (need to accept the Airbus license) - The password will be provided in the confirmation email. ---------------------------------------------------------------------------- () The PAirMax dataset is a collection of data with the aim of assessing the performance of pansharpening algorithms. The data collection includes 5 test cases selected at full resolution (FR), acquired by two sensors belonging to the Airbus' constellation of high-resolution imaging satellites. Moreover, 5 related test cases at reduced resolution (RR), simulated according to the Wald’s protocol, are included, thus resulting in 10 challenging test cases for pansharpening performance assessment. For further details, please, refer to the paper: [1] G. Vivone, M. Dalla Mura, A. Garzelli, and F. Pacifici, "A Benchmarking Protocol for Pansharpening: Dataset, Pre-processing, and Quality Assessment," IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2021.
Authors
- Dalla Mura, Mauro
This database called "TMD-CAPTIMOVE" provides transportation mode labelled data collected by 34 volunteers for a total time duration of around 48 hours. Considered transportation modes are: On-foot (Walking, Stairs, Elevators), Bike, Scooter, Bus, Tram. The number of labels is 11: Still, Walk, Upstairs, Downstairs, Elevator up, Elevator down, Bike, Electric scooter, kick scooter, Bus, Tram. Sensor data are: Acceleration (m/s²), angular rate (°/s), atmospheric pressure (hPa), heart rate (beat per minute (BPM)). The sampling frequency for all data is 32 Hertz.
Authors
- Fourati, Hassen
This dataset contains all data, source code, pre-trained computational predictive models and experimental results related to: Hueber T., Tatulli E., Girin L., Schwatz, J-L "How predictive can be predictions in the neurocognitive processing of auditory and audiovisual speech? A deep learning study." (biorXiv preprint https://doi.org/10.1101/471581). Raw data are extracted from the publicly available database NTCD-TIMIT (10.5281/zenodo.260228). Audio recordings are available in the audio_clean/ directoryPost-processed lip image sequences are available in the lips_roi/ directory (67x67 pixels, 8bits, obtained by lossless inverse DCT-2D transform from the DCT feature available in the original repository of NTCD-TIMIT)Phonetic segmentation (extracted from NTCD-TIMIT original zenodo repository) is available in the HTK MLF file volunteer_labelfiles.mlfAudio features (MFCC-spectrogram and log-spectrogram) are available in the mfcc_16k/ and fft_16k/ directories. Models (audio-only, video-only and audiovisual, based on deep feed-forward neural networks and/or convolutional neural network, in .h5 format, trained with Keras 2.0 toolkit) and data normalization parameters (in .dat scikit-learn format) are available in models_mfcc/ and models_logspectro/ directoriesPredicted and target (ground truth) MFCC-spectro (resp. log-spectro) for the test databases (1909 sentences), and for the different values of (\tau_p) or (\tau_f) are available in pred_testdb_mfccspectro/ (resp. pred_testdb_logspectro/) directorySource code for extracting audio features, training and evaluating the models is available on GitHub https://github.com/thueber/DeepPredSpeech/All directories have been zipped before upload.Feel free to contact me for more details.Thomas Hueber, Ph. D., CNRS research fellow, GIPSA-lab, Grenoble, France, [email protected]
Authors
- Hueber, Thomas ;
- Tatulli, Eric ;
- Girin, Laurent ;
- Schwartz, Jean-Luc
This dataset contains all data, source code, pre-trained computational predictive models and experimental results related to: Hueber T., Tatulli E., Girin L., Schwatz, J-L "How predictive can be predictions in the neurocognitive processing of auditory and audiovisual speech? A deep learning study." (biorXiv preprint https://doi.org/10.1101/471581). Raw data are extracted from the publicly available database NTCD-TIMIT (10.5281/zenodo.260228). Audio recordings are available in the audio_clean/ directory Post-processed lip image sequences are available in the lips_roi/ directory (67x67 pixels, 8bits, obtained by lossless inverse DCT-2D transform from the DCT feature available in the original repository of NTCD-TIMIT) Phonetic segmentation (extracted from NTCD-TIMIT original zenodo repository) is available in the HTK MLF file volunteer_labelfiles.mlf Audio features (MFCC-spectrogram and log-spectrogram) are available in the mfcc_16k/ and fft_16k/ directories. Models (audio-only, video-only and audiovisual, based on deep feed-forward neural networks and/or convolutional neural network, in .h5 format, trained with Keras 2.0 toolkit) and data normalization parameters (in .dat scikit-learn format) are available in models_mfcc/ and models_logspectro/ directories Predicted and target (ground truth) MFCC-spectro (resp. log-spectro) for the test databases (1909 sentences), and for the different values of (\tau_p) or (\tau_f) are available in pred_testdb_mfccspectro/ (resp. pred_testdb_logspectro/) directory Source code for extracting audio features, training and evaluating the models is available on GitHub https://github.com/thueber/DeepPredSpeech/ All directories have been zipped before upload. Feel free to contact me for more details. Thomas Hueber, Ph. D., CNRS research fellow, GIPSA-lab, Grenoble, France, [email protected]
Authors
- Hueber, Thomas ;
- Tatulli, Eric ;
- Girin, Laurent ;
- Schwartz, Jean-Luc
This dataset contains all data, source code, pre-trained computational predictive models and experimental results related to: Hueber T., Tatulli E., Girin L., Schwatz, J-L "How predictive can be predictions in the neurocognitive processing of auditory and audiovisual speech? A deep learning study." (biorXiv preprint https://doi.org/10.1101/471581). Raw data are extracted from the publicly available database NTCD-TIMIT (10.5281/zenodo.260228). Audio recordings are available in the audio_clean/ directory Post-processed lip image sequences are available in the lips_roi/ directory (67x67 pixels, 8bits, obtained by lossless inverse DCT-2D transform from the DCT feature available in the original repository of NTCD-TIMIT) Phonetic segmentation (extracted from NTCD-TIMIT original zenodo repository) is available in the HTK MLF file volunteer_labelfiles.mlf Audio features (MFCC-spectrogram and log-spectrogram) are available in the mfcc_16k/ and fft_16k/ directories. Models (audio-only, video-only and audiovisual, based on deep feed-forward neural networks and/or convolutional neural network, in .h5 format, trained with Keras 2.0 toolkit) and data normalization parameters (in .dat scikit-learn format) are available in models_mfcc/ and models_logspectro/ directories Predicted and target (ground truth) MFCC-spectro (resp. log-spectro) for the test databases (1909 sentences), and for the different values of (\tau_p) or (\tau_f) are available in pred_testdb_mfccspectro/ (resp. pred_testdb_logspectro/) directory Source code for extracting audio features, training and evaluating the models is available on GitHub https://github.com/thueber/DeepPredSpeech/ All directories have been zipped before upload. Feel free to contact me for more details. Thomas Hueber, Ph. D., CNRS research fellow, GIPSA-lab, Grenoble, France, [email protected]
Authors
- Hueber, Thomas ;
- Tatulli, Eric ;
- Girin, Laurent ;
- Schwartz, Jean-Luc
CSF18 - Multimodal database of French Cued-speech (revised in 2022) Dataset used in "Visual recognition of continuous Cued Speech using a tandem CNN-HMM approach", by Liu, Hueber, Feng, Beautemps (submitted to Interspeech 2018) 476 sentences (i.e. 2 repetitions of 238 sentences) uttered by a professional French Cued-speech coder video/ PNG images, 576x720, 50fps (after deinterleave) audio/ WAV, 16kHz, 16bits prompt.txt: Text prompt of the recorded sentences corpus_mlf.txt: Phonetic transcription aligned on the audio signal (HTK format, Master Label File) obtained using the LiaPhon phonetizer and a forced-alignment HMM-based procedure (no manual check) corpus_mlf_updated_icassp2022.txt: Manually checked/cleaned version of corpus_mlf.txt (see Sankar et al., ICASSP 2022 paper) phonelist.txt: list of the 34 labels used to encode French phonemes at GIPSA-lab.
Authors
- Liu, Li ;
- Hueber, Thomas ;
- Feng, Gang ;
- Beautemps, Denis ;
- Sankar, Sanjana
CSF18 - Multimodal database of French Cued-speech (revised in 2022) Dataset used in "Visual recognition of continuous Cued Speech using a tandem CNN-HMM approach", by Liu, Hueber, Feng, Beautemps (submitted to Interspeech 2018) 476 sentences (i.e. 2 repetitions of 238 sentences) uttered by a professional French Cued-speech coder video/ PNG images, 576x720, 50fps (after deinterleave) audio/ WAV, 16kHz, 16bits prompt.txt: Text prompt of the recorded sentences corpus_mlf.txt: Phonetic transcription aligned on the audio signal (HTK format, Master Label File) obtained using the LiaPhon phonetizer and a forced-alignment HMM-based procedure (no manual check) corpus_mlf_updated_icassp2022.txt: Manually checked/cleaned version of corpus_mlf.txt (see Sankar et al., ICASSP 2022 paper) phonelist.txt: list of the 34 labels used to encode French phonemes at GIPSA-lab.
Authors
- Liu, Li ;
- Hueber, Thomas ;
- Feng, Gang ;
- Beautemps, Denis ;
- Sankar, Sanjana
CSF18 - Multimodal database of French Cued-speech Dataset used in "Visual recognition of continuous Cued Speech using a tandem CNN-HMM approach", by Liu, Hueber, Feng, Beautemps (submitted to Interspeech 2018)476 sentences (i.e. 2 repetitions of 238 sentences) uttered by a professional French Cued-speech codervideo/ PNG images, 576x720, 50fps (after deinterleave)audio/ WAV, 16kHz, 16bitsprompt.txt: Text prompt of the recorded sentences corpus_mlf.txt: Phonetic transcription aligned on the audio signal (HTK format, Master Label File) obtained using the LiaPhon phonetizer and a forced-alignment HMM-based procedure (no manual check)phonelist.txt: list of the 34 labels used to encode French phonemes at GIPSA-lab.
Authors
- Liu, Li ;
- Hueber, Thomas ;
- Feng, Gang ;
- Beautemps, Denis