“ASR:2015-06-15”版本间的差异

2015年6月25日 (四) 02:46的最后版本

Speech Processing

AM development

Environment

grid-14 does not work --mengyuan
grid-15 runs slowly

RNN AM

morpheme RNN --zhiyuan
RNN MPE --zhiyuan and xuewei

Mic-Array

hold
compute EER with kaldi

RNN-DAE(Deep based Auto-Encode-RNN)

hold
deliver to mengyuan

http://cslt.riit.tsinghua.edu.cn/cgi-bin/cvss/cvss_request.pl?account=zhangzy&step=view_request&cvssid=261

Speaker ID

DNN-based sid --Lantian

http://cslt.riit.tsinghua.edu.cn/cgi-bin/cvss/cvss_request.pl?account=zhangzy&step=view_request&cvssid=327

Ivector&Dvector based ASR

hold --Tian Lan
Cluster the speakers to speaker-classes, then using the distance or the posterior-probability as the metric
dark-konowlege using i-vector
train on wsj(testbase dev93+evl92)

--hold

Dark knowledge

test random last output layer when train MPE--zhiyuan

language vector

hold --xuewei
train using chinese and chiglish

Text Processing

RNN LM

character-lm rnn(hold)
lstm+rnn

check the lstm-rnnlm code about how to Initialize and update learning rate.(hold)

W2V based document classification

APSIPA paper
CNN adapt to resolve the low resource problem

Pair-wise LM

draft paper of journal

Order representation

modify the objective function(hold)
sup-sampling method to solve the low frequence word(hold)
journal paper

binary vector

nips paper

Stochastic ListNet

done

relation classifier

done

plan to do

combine LDA with neural network

@@ 第3行： / 第3行： @@
 ==== Environment ====
-*
+*grid-14 does not work --mengyuan
+*grid-15 runs slowly
 ==== RNN AM====
-*morpheme RNN-zhiyuan
+*morpheme RNN --zhiyuan
+*RNN MPE --zhiyuan and xuewei
-* ==== Mic-Array ====
+==== Mic-Array ====
 * hold
-* Change the prediction from  fbank to spectrum features
-* investigate alpha parameter in time domian and frquency domain
-* ALPHA>=0, using data generated by reverber toolkit
-* consider theta
 * compute EER with kaldi
@@ 第21行： / 第20行： @@
 ===Speaker ID===
-*  DNN-based sid --Tian Lan
+*  DNN-based sid --Lantian
 :* http://cslt.riit.tsinghua.edu.cn/cgi-bin/cvss/cvss_request.pl?account=zhangzy&step=view_request&cvssid=327
 ===Ivector&Dvector based ASR===
-*  hold --Tian Lan
+* hold --Tian Lan
 * Cluster the speakers to speaker-classes, then using the distance or the posterior-probability as the metric
-* Direct using the dark-knowledge strategy to do the ivector training.
+* dark-konowlege using i-vector
-:* http://cslt.riit.tsinghua.edu.cn/cgi-bin/cvss/cvss_request.pl?step=view_request&cvssid=340
-* Ivector dimention is smaller, performance is better
-* Augument to hidden layer is better than input layer
 * train on wsj(testbase dev93+evl92)
+:*--hold
 ===Dark knowledge===
-* Ensemble using 100h dataset to construct diffrernt structures -- Mengyuan
-:*http://cslt.riit.tsinghua.edu.cn/cgi-bin/cvss/cvss_request.pl?account=zxw&step=view_request&cvssid=264 --Zhiyong Zhang
-* adaptation English and Chinglish
-:* Try to improve the chinglish performance extremly
-* unsupervised training with wsj contributes to aurora4 model --Xiangyu Zeng
-* test large database with AMIDA
-* test hidden layer knowledge transfer--xuewei
 * test random last output layer when train MPE--zhiyuan
-===bilingual recognition===
-* hold
-:* http://cslt.riit.tsinghua.edu.cn/cgi-bin/cvss/cvss_request.pl?account=zxw&step=view_request&cvssid=359 --Zhiyuan Tang and Mengyuan
 ===language vector===
-* train DNN with language vector--xuewei
+* hold --xuewei
+* train using chinese and chiglish
 ==Text Processing==

“ASR:2015-06-15”版本间的差异

2015年6月25日 (四) 02:46的最后版本

目录

Speech Processing

AM development

Environment

RNN AM

Mic-Array

RNN-DAE(Deep based Auto-Encode-RNN)

Speaker ID

Ivector&Dvector based ASR

Dark knowledge

language vector

Text Processing

RNN LM

W2V based document classification

Pair-wise LM

Order representation

binary vector

Stochastic ListNet

relation classifier

plan to do

导航菜单

个人工具

名字空间

变种

查看

操作

搜索

导航

工具