diff --git a/.gitattributes b/.gitattributes
new file mode 100644
index 00000000..7c32d5f7
--- /dev/null
+++ b/.gitattributes
@@ -0,0 +1 @@
+*.jar filter=lfs diff=lfs merge=lfs -text
diff --git a/LICENSE b/LICENSE
deleted file mode 100644
index 7bf9011a..00000000
--- a/LICENSE
+++ /dev/null
@@ -1,201 +0,0 @@
- Apache License
- Version 2.0, January 2004
- http://www.apache.org/licenses/
-
- TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
-
- 1. Definitions.
-
- "License" shall mean the terms and conditions for use, reproduction,
- and distribution as defined by Sections 1 through 9 of this document.
-
- "Licensor" shall mean the copyright owner or entity authorized by
- the copyright owner that is granting the License.
-
- "Legal Entity" shall mean the union of the acting entity and all
- other entities that control, are controlled by, or are under common
- control with that entity. For the purposes of this definition,
- "control" means (i) the power, direct or indirect, to cause the
- direction or management of such entity, whether by contract or
- otherwise, or (ii) ownership of fifty percent (50%) or more of the
- outstanding shares, or (iii) beneficial ownership of such entity.
-
- "You" (or "Your") shall mean an individual or Legal Entity
- exercising permissions granted by this License.
-
- "Source" form shall mean the preferred form for making modifications,
- including but not limited to software source code, documentation
- source, and configuration files.
-
- "Object" form shall mean any form resulting from mechanical
- transformation or translation of a Source form, including but
- not limited to compiled object code, generated documentation,
- and conversions to other media types.
-
- "Work" shall mean the work of authorship, whether in Source or
- Object form, made available under the License, as indicated by a
- copyright notice that is included in or attached to the work
- (an example is provided in the Appendix below).
-
- "Derivative Works" shall mean any work, whether in Source or Object
- form, that is based on (or derived from) the Work and for which the
- editorial revisions, annotations, elaborations, or other modifications
- represent, as a whole, an original work of authorship. For the purposes
- of this License, Derivative Works shall not include works that remain
- separable from, or merely link (or bind by name) to the interfaces of,
- the Work and Derivative Works thereof.
-
- "Contribution" shall mean any work of authorship, including
- the original version of the Work and any modifications or additions
- to that Work or Derivative Works thereof, that is intentionally
- submitted to Licensor for inclusion in the Work by the copyright owner
- or by an individual or Legal Entity authorized to submit on behalf of
- the copyright owner. For the purposes of this definition, "submitted"
- means any form of electronic, verbal, or written communication sent
- to the Licensor or its representatives, including but not limited to
- communication on electronic mailing lists, source code control systems,
- and issue tracking systems that are managed by, or on behalf of, the
- Licensor for the purpose of discussing and improving the Work, but
- excluding communication that is conspicuously marked or otherwise
- designated in writing by the copyright owner as "Not a Contribution."
-
- "Contributor" shall mean Licensor and any individual or Legal Entity
- on behalf of whom a Contribution has been received by Licensor and
- subsequently incorporated within the Work.
-
- 2. Grant of Copyright License. Subject to the terms and conditions of
- this License, each Contributor hereby grants to You a perpetual,
- worldwide, non-exclusive, no-charge, royalty-free, irrevocable
- copyright license to reproduce, prepare Derivative Works of,
- publicly display, publicly perform, sublicense, and distribute the
- Work and such Derivative Works in Source or Object form.
-
- 3. Grant of Patent License. Subject to the terms and conditions of
- this License, each Contributor hereby grants to You a perpetual,
- worldwide, non-exclusive, no-charge, royalty-free, irrevocable
- (except as stated in this section) patent license to make, have made,
- use, offer to sell, sell, import, and otherwise transfer the Work,
- where such license applies only to those patent claims licensable
- by such Contributor that are necessarily infringed by their
- Contribution(s) alone or by combination of their Contribution(s)
- with the Work to which such Contribution(s) was submitted. If You
- institute patent litigation against any entity (including a
- cross-claim or counterclaim in a lawsuit) alleging that the Work
- or a Contribution incorporated within the Work constitutes direct
- or contributory patent infringement, then any patent licenses
- granted to You under this License for that Work shall terminate
- as of the date such litigation is filed.
-
- 4. Redistribution. You may reproduce and distribute copies of the
- Work or Derivative Works thereof in any medium, with or without
- modifications, and in Source or Object form, provided that You
- meet the following conditions:
-
- (a) You must give any other recipients of the Work or
- Derivative Works a copy of this License; and
-
- (b) You must cause any modified files to carry prominent notices
- stating that You changed the files; and
-
- (c) You must retain, in the Source form of any Derivative Works
- that You distribute, all copyright, patent, trademark, and
- attribution notices from the Source form of the Work,
- excluding those notices that do not pertain to any part of
- the Derivative Works; and
-
- (d) If the Work includes a "NOTICE" text file as part of its
- distribution, then any Derivative Works that You distribute must
- include a readable copy of the attribution notices contained
- within such NOTICE file, excluding those notices that do not
- pertain to any part of the Derivative Works, in at least one
- of the following places: within a NOTICE text file distributed
- as part of the Derivative Works; within the Source form or
- documentation, if provided along with the Derivative Works; or,
- within a display generated by the Derivative Works, if and
- wherever such third-party notices normally appear. The contents
- of the NOTICE file are for informational purposes only and
- do not modify the License. You may add Your own attribution
- notices within Derivative Works that You distribute, alongside
- or as an addendum to the NOTICE text from the Work, provided
- that such additional attribution notices cannot be construed
- as modifying the License.
-
- You may add Your own copyright statement to Your modifications and
- may provide additional or different license terms and conditions
- for use, reproduction, or distribution of Your modifications, or
- for any such Derivative Works as a whole, provided Your use,
- reproduction, and distribution of the Work otherwise complies with
- the conditions stated in this License.
-
- 5. Submission of Contributions. Unless You explicitly state otherwise,
- any Contribution intentionally submitted for inclusion in the Work
- by You to the Licensor shall be under the terms and conditions of
- this License, without any additional terms or conditions.
- Notwithstanding the above, nothing herein shall supersede or modify
- the terms of any separate license agreement you may have executed
- with Licensor regarding such Contributions.
-
- 6. Trademarks. This License does not grant permission to use the trade
- names, trademarks, service marks, or product names of the Licensor,
- except as required for reasonable and customary use in describing the
- origin of the Work and reproducing the content of the NOTICE file.
-
- 7. Disclaimer of Warranty. Unless required by applicable law or
- agreed to in writing, Licensor provides the Work (and each
- Contributor provides its Contributions) on an "AS IS" BASIS,
- WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
- implied, including, without limitation, any warranties or conditions
- of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
- PARTICULAR PURPOSE. You are solely responsible for determining the
- appropriateness of using or redistributing the Work and assume any
- risks associated with Your exercise of permissions under this License.
-
- 8. Limitation of Liability. In no event and under no legal theory,
- whether in tort (including negligence), contract, or otherwise,
- unless required by applicable law (such as deliberate and grossly
- negligent acts) or agreed to in writing, shall any Contributor be
- liable to You for damages, including any direct, indirect, special,
- incidental, or consequential damages of any character arising as a
- result of this License or out of the use or inability to use the
- Work (including but not limited to damages for loss of goodwill,
- work stoppage, computer failure or malfunction, or any and all
- other commercial damages or losses), even if such Contributor
- has been advised of the possibility of such damages.
-
- 9. Accepting Warranty or Additional Liability. While redistributing
- the Work or Derivative Works thereof, You may choose to offer,
- and charge a fee for, acceptance of support, warranty, indemnity,
- or other liability obligations and/or rights consistent with this
- License. However, in accepting such obligations, You may act only
- on Your own behalf and on Your sole responsibility, not on behalf
- of any other Contributor, and only if You agree to indemnify,
- defend, and hold each Contributor harmless for any liability
- incurred by, or claims asserted against, such Contributor by reason
- of your accepting any such warranty or additional liability.
-
- END OF TERMS AND CONDITIONS
-
- APPENDIX: How to apply the Apache License to your work.
-
- To apply the Apache License to your work, attach the following
- boilerplate notice, with the fields enclosed by brackets "[]"
- replaced with your own identifying information. (Don't include
- the brackets!) The text should be enclosed in the appropriate
- comment syntax for the file format. We also recommend that a
- file or class name and description of purpose be included on the
- same "printed page" as the copyright notice for easier
- identification within third-party archives.
-
- Copyright 2017, The TensorFlow Authors.
-
- Licensed under the Apache License, Version 2.0 (the "License");
- you may not use this file except in compliance with the License.
- You may obtain a copy of the License at
-
- http://www.apache.org/licenses/LICENSE-2.0
-
- Unless required by applicable law or agreed to in writing, software
- distributed under the License is distributed on an "AS IS" BASIS,
- WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
- See the License for the specific language governing permissions and
- limitations under the License.
diff --git a/PUBLICATIONS.md b/PUBLICATIONS.md
deleted file mode 100644
index f8ac7eed..00000000
--- a/PUBLICATIONS.md
+++ /dev/null
@@ -1,1114 +0,0 @@
-# List of publications using Lingvo.
-
-
-
-## Translation
-
-
-
-
-
-
-
-|
-[1]
- |
-
-Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun,
- Y. Cao, Q. Gao, K. Macherey, J. Klingner, A. Shah, M. Johnson, X. Liu,
- L. Kaiser, S. Gouws, Y. Kato, T. Kudo, H. Kazawa, K. Stevens, G. Kurian,
- N. Patil, W. Wang, C. Young, J. Smith, J. Riesa, A. Rudnick, O. Vinyals,
- G. Corrado, M. Hughes, and J. Dean, “Google's neural machine translation
- system: Bridging the gap between human and machine translation,” tech. rep.,
- 2016.
-[ pdf ]
-
- |
-
-
-
-
-|
-[2]
- |
-
-M. Johnson, M. Schuster, Q. V. Le, M. Krikun, Y. Wu, Z. Chen, N. Thorat,
- F. Viégas, M. Wattenberg, G. Corrado, M. Hughes, and J. Dean,
- “Google's multilingual neural machine translation system: Enabling
- zero-shot translation,” Transactions of the Association for
- Computational Linguistics, vol. 5, pp. 339--351, 2017.
-[ DOI |
-pdf ]
-
- |
-
-
-
-
-|
-[3]
- |
-
-A. Eriguchi, M. Johnson, O. Firat, H. Kazawa, and W. Macherey, “Zero-shot
- cross-lingual classification using multilingual neural machine translation,”
- arXiv preprint arXiv:1809.04686, 2018.
-[ pdf ]
-
- |
-
-
-
-
-|
-[4]
- |
-
-A. Bapna, M. X. Chen, O. Firat, Y. Cao, and Y. Wu, “Training deeper neural
- machine translation models with transparent attention,” in Proc.
- Conference on Empirical Methods in Natural Language Processing (EMNLP),
- 2018.
-[ pdf ]
-
- |
-
-
-
-
-|
-[5]
- |
-
-C. Cherry, G. Foster, A. Bapna, O. Firat, and W. Macherey, “Revisiting
- character-based neural machine translation with capacity and compression,”
- in Proc. Conference on Empirical Methods in Natural Language Processing
- (EMNLP), 2018.
-[ pdf ]
-
- |
-
-
-
-
-|
-[6]
- |
-
-M. X. Chen, O. Firat, A. Bapna, M. Johnson, W. Macherey, G. Foster, L. Jones,
- M. Schuster, N. Shazeer, N. Parmar, A. Vaswani, J. Uszkoreit, L. Kaiser,
- Z. Chen, Y. Wu, and M. Hughes, “The Best of Both Worlds: Combining Recent
- Advances in Neural Machine Translation,” in Proc. Annual Meeting of
- the Association for Computational Linguistics (ACL), 2018.
-[ pdf ]
-
- |
-
-
-
-
-|
-[7]
- |
-
-J. Kuczmarski and M. Johnson, “Gender-aware natural language translation,”
- 2018.
-[ pdf ]
-
- |
-
-
-
-
-|
-[8]
- |
-
-R. Aharoni, M. Johnson, and O. Firat, “Massively multilingual neural machine
- translation,” 2019.
-[ pdf ]
-
- |
-
-
-
-
-|
-[9]
- |
-
-J. Luo, Y. Cao, and R. Barzilay, “Neural decipherment via minimum-cost flow:
- From ugaritic to linear b,” 2019.
-[ http ]
-
- |
-
-
-
-
-|
-[10]
- |
-
-N. Arivazhagan, C. Cherry, W. Macherey, C.-C. Chiu, S. Yavuz, R. Pang, W. Li,
- and C. Raffel, “Monotonic infinite lookback attention for simultaneous
- machine translation,” in Proc. Annual Meeting of the Association for
- Computational Linguistics (ACL), 2019.
-[ http ]
-
- |
-
-
-
-
-|
-[11]
- |
-
-M. Freitag, I. Caswell, and S. Roy, “Ape at scale and its implications on mt
- evaluation biases,” 2019.
-[ pdf |
-http ]
-
- |
-
-
-
-
-|
-[12]
- |
-
-N. Arivazhagan, A. Bapna, O. Firat, D. Lepikhin, M. Johnson, M. Krikun, M. X.
- Chen, Y. Cao, G. Foster, C. Cherry, W. Macherey, Z. Chen, and Y. Wu,
- “Massively multilingual neural machine translation in the wild: Findings and
- challenges,” 2019.
-[ arXiv |
-http ]
-
- |
-
-
-
-
-|
-[13]
- |
-
-Y. Huang, Y. Cheng, A. Bapna, O. Firat, M. X. Chen, D. Chen, H. Lee, J. Ngiam,
- Q. V. Le, Y. Wu, and Z. Chen, “Gpipe: Efficient training of giant neural
- networks using pipeline parallelism,” in Advances in Neural Information
- Processing Systems, 2019.
-[ http ]
-
- |
-
-
-
-## Speech recognition
-
-
-
-
-
-
-
-|
-[1]
- |
-
-C.-C.Chiu, T. N. Sainath, Y. Wu, R. Prabhavalkar, P. Nguyen, Z. Chen,
- A. Kannan, R. J. Weiss, K. Rao, K. Gonina, N. Jaitly, B. Li, J. Chorowski,
- and M. Bacchiani, “State-of-the-art speech recognition with
- sequence-to-sequence models,” in Proc. IEEE International Conference
- on Acoustics, Speech, and Signal Processing (ICASSP), 2018.
-[ pdf ]
-
- |
-
-
-
-
-|
-[2]
- |
-
-S. Toshniwal, T. N. Sainath, R. J. Weiss, B. Li, P. Moreno, E. Weinstein, and
- K. Rao, “Multilingual speech recognition with a single end-to-end model,”
- in Proc. IEEE International Conference on Acoustics, Speech, and
- Signal Processing (ICASSP), 2018.
-[ pdf ]
-
- |
-
-
-
-
-|
-[3]
- |
-
-B. Li, T. N. Sainath, K. Sim, M. Bacchiani, E. Weinstein, P. Nguyen, Z. Chen,
- Y. Wu, and K. Rao, “Multi-Dialect Speech Recognition With a Single
- Sequence-to-Sequence Model,” in Proc. IEEE International Conference
- on Acoustics, Speech, and Signal Processing (ICASSP), 2018.
-[ pdf ]
-
- |
-
-
-
-
-|
-[4]
- |
-
-T. N. Sainath, P. Prabhavalkar, S. Kumar, S. Lee, A. Kannan, D. Rybach,
- V. Schogol, P. Nguyen, B. Li, Y. Wu, Z. Chen, and C. C. Chiu, “No Need for
- a Lexicon? Evaluating the Value of the Pronunciation Lexica in End-to-End
- Models,” in Proc. IEEE International Conference on Acoustics,
- Speech, and Signal Processing (ICASSP), 2018.
-[ pdf ]
-
- |
-
-
-
-
-|
-[5]
- |
-
-D. Lawson, C. C. Chiu, G. Tucker, C. Raffel, K. Swersky, and N. Jaitly,
- “Learning hard alignments with variational inference,” in Proc.
- IEEE International Conference on Acoustics, Speech, and Signal Processing
- (ICASSP), 2018.
-[ pdf ]
-
- |
-
-
-
-
-|
-[6]
- |
-
-A. Kannan, Y. Wu, P. Nguyen, T. N. Sainath, Z. Chen, and R. Prabhavalkar, “An
- analysis of incorporating an external language model into a
- sequence-to-sequence model,” in Proc. IEEE International Conference
- on Acoustics, Speech, and Signal Processing (ICASSP), 2018.
-[ pdf ]
-
- |
-
-
-
-
-|
-[7]
- |
-
-R. Prabhavalkar, T. N. Sainath, Y. Wu, P. Nguyen, Z. Chen, C. C. Chiu, and
- A. Kannan, “Minimum Word Error Rate Training for Attention-based
- Sequence-to-sequence Models,” in Proc. IEEE International Conference
- on Acoustics, Speech, and Signal Processing (ICASSP), 2018.
-[ pdf ]
-
- |
-
-
-
-
-|
-[8]
- |
-
-T. N. Sainath, C. C. Chiu, R. Prabhavalkar, A. Kannan, Y. Wu, P. Nguyen, and
- Z. C. Z, “Improving the Performance of Online Neural Transducer Models,”
- in Proc. IEEE International Conference on Acoustics, Speech, and
- Signal Processing (ICASSP), 2018.
-[ pdf ]
-
- |
-
-
-
-
-|
-[9]
- |
-
-C. C. Chiu and C. Raffel, “Monotonic Chunkwise Attention,” in Proc.
- International Conference on Learning Representations (ICLR), 2018.
-[ pdf ]
-
- |
-
-
-
-
-|
-[10]
- |
-
-I. Williams, A. Kannan, P. Aleksic, D. Rybach, and T. N. S. TN, “Contextual
- Speech Recognition in End-to-End Neural Network Systems using Beam Search,”
- in Proc. Interspeech, 2018.
-[ pdf ]
-
- |
-
-
-
-
-|
-[11]
- |
-
-C. C. Chiu, A. Tripathi, K. Chou, C. Co, N. Jaitly, D. Jaunzeikare, A. Kannan,
- P. Nguyen, H. Sak, A. Sankar, J. Tansuwan, N. Wan, Y. Wu, and X. Zhang,
- “Speech recognition for medical conversations,” in Proc.
- Interspeech, 2018.
-[ pdf ]
-
- |
-
-
-
-
-|
-[12]
- |
-
-R. Pang, T. N. Sainath, R. Prabhavalkar, S. Gupta, Y. Wu, S. Zhang, and C. C.
- Chiu, “Compression of End-to-End Models,” in Proc. Interspeech,
- 2018.
-[ pdf ]
-
- |
-
-
-
-
-|
-[13]
- |
-
-S. Toshniwal, A. Kannan, C. C. Chiu, Y. Wu, T. N. Sainath, and K. Livescu, “A
- comparison of techniques for language model integration in encoder-decoder
- speech recognition,” in Proc. IEEE Spoken Language Technology
- Workshop (SLT), 2018.
-[ pdf ]
-
- |
-
-
-
-
-|
-[14]
- |
-
-G. Pundak, T. N. Sainath, R. Prabhavalkar, A. Kannan, and D. Zhao, “Deep
- context: End-to-end contextual speech recognition,” in Proc. IEEE
- Spoken Language Technology Workshop (SLT), 2018.
-[ pdf ]
-
- |
-
-
-
-
-|
-[15]
- |
-
-B. Li, Y. Zhang, T. N. Sainath, Y. Wu, and W. Chan, “Bytes are all you need:
- End-to-end multilingual speech recognition and synthesis with bytes,” in
- Proc. IEEE International Conference on Acoustics, Speech, and Signal
- Processing (ICASSP), 2019.
-[ pdf ]
-
- |
-
-
-
-
-|
-[16]
- |
-
-J. Guo, T. N. Sainath, and R. J. Weiss, “A spelling correction model for
- end-to-end speech recognition,” in Proc. IEEE International
- Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2019.
-[ pdf ]
-
- |
-
-
-
-
-|
-[17]
- |
-
-U. Alon, G. Pundak, and T. N. Sainath, “Contextual speech recognition with
- difficult negative training examples,” in Proc. IEEE International
- Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2019.
-[ pdf ]
-
- |
-
-
-
-
-|
-[18]
- |
-
-Y. Qin, N. Carlini, I. Goodfellow, G. Cottrell, and C. Raffel, “Imperceptible,
- robust, and targeted adversarial examples for automatic speech recognition,”
- in Proc. International Conference on Machine Learning (ICML), 2019.
-[ pdf ]
-
- |
-
-
-
-
-|
-[19]
- |
-
-D. S. Park, W. Chan, Y. Zhang, C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le,
- “SpecAugment: A Simple Data Augmentation Method for Automatic Speech
- Recognition,” in arXiv, 2019.
-[ pdf ]
-
- |
-
-
-
-
-|
-[20]
- |
-
-B. Li, T. N. Sainath, R. Pang, and Z. Wu, “Semi-supervised training for
- end-to-end models via weak distillation,” in Proc. IEEE International
- Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2019.
-[ pdf ]
-
- |
-
-
-
-
-|
-[21]
- |
-
-S.-Y. Chang, R. Prabhavalkar, Y. He, T. N. Sainath, and G. Simko, “Joint
- endpointing and decoding with end-to-end models,” in Proc. IEEE
- International Conference on Acoustics, Speech, and Signal Processing
- (ICASSP), 2019.
-[ pdf ]
-
- |
-
-
-
-
-|
-[22]
- |
-
-J. Heymann, K. C. Sim, and B. Li, “Improving ctc using stimulated learning for
- sequence modeling,” in Proc. IEEE International Conference on
- Acoustics, Speech, and Signal Processing (ICASSP), 2019.
-[ pdf ]
-
- |
-
-
-
-
-|
-[23]
- |
-
-A. Bruguier, R. Prabhavalkar, G. Pundak, and T. N. Sainath, “Phoebe:
- Pronunciation-aware contextualization for end-to-end speech recognition,” in
- Proc. IEEE International Conference on Acoustics, Speech, and Signal
- Processing (ICASSP), 2019.
-[ pdf ]
-
- |
-
-
-
-
-|
-[24]
- |
-
-Y. He, T. N. Sainath, R. Prabhavalkar, I. McGraw, R. Alvarez, D. Zhao,
- D. Rybach, A. Kannan, Y. Wu, R. Pang, Q. Liang, D. Bhatia, Y. Shangguan,
- B. Li, G. Pundak, K. C. Sim, T. Bagby, S.-Y. Chang, K. Rao, and
- A. Gruenstein, “Streaming end-to-end speech recognition for mobile
- devices,” in Proc. IEEE International Conference on Acoustics,
- Speech, and Signal Processing (ICASSP), 2019.
-[ pdf ]
-
- |
-
-
-
-
-|
-[25]
- |
-
-K. Irie, R. Prabhavalkar, A. Kannan, A. Bruguier, D. Rybach, and P. Nguyen,
- “On the choice of modeling unit for sequence-to-sequence speech
- recognition,” in Proc. Interspeech, 2019.
-[ pdf ]
-
- |
-
-
-
-
-|
-[26]
- |
-
-C. Peyser, H. Zhang, T. N. Sainath, and Z. Wu, “Improving Performance of
- End-to-End ASR on Numeric Sequences,” in Proc. Interspeech, 2019.
-[ pdf ]
-
- |
-
-
-
-
-|
-[27]
- |
-
-D. Zhao, T. N. Sainath, D. Rybach, D. Bhatia, B. Li, and R. Pang,
- “Shallow-fusion end-to-end contextual biasing,” in Proc. Interspeech,
- 2019.
-[ pdf ]
-
- |
-
-
-
-
-|
-[28]
- |
-
-T. N. Sainath, R. Pang, D. Rybach, Y. He, R. Prabhavalkar, W. Li, M. Visontai,
- Q. Liang, T. Strohman, Y. Wu, I. McGraw, and C.-C. Chiu, “Two-pass
- end-to-end speech recognition,” in Proc. Interspeech, 2019.
-[ pdf ]
-
- |
-
-
-
-
-|
-[29]
- |
-
-C.-C. Chiu, W. Han, Y. Zhang, R. Pang, S. Kishchenko, P. Nguyen, A. Narayanan,
- H. Liao, S. Zhang, A. Kannan, R. Prabhavalkar, Z. Chen, T. Sainath, and
- Y. Wu, “A comparison of end-to-end models for long-form speech
- recognition,” 2019.
-[ pdf ]
-
- |
-
-
-
-
-|
-[30]
- |
-
-A. Narayanan, R. Prabhavalkar, C. Chiu, D. Rybach, T. Sainath, and T. Strohman,
- “Recognizing long-form speech using streaming end-to-end models,” 2019.
-[ pdf ]
-
- |
-
-
-
-
-|
-[31]
- |
-
-T. N. Sainath, R. Pang, R. Weiss, Y. He, C.-C. Chiu, and T. Strohman, “An
- attention-based joint acoustic and text on-device end-to-end model,” in
- Proc. IEEE International Conference on Acoustics, Speech, and Signal
- Processing (ICASSP), 2020.
-
- |
-
-
-
-
-|
-[32]
- |
-
-Z. Lu, L. Cao, Y. Zhang, C.-C. Chiu, and J. Fan, “Speech sentiment analysis
- via pre-trained features from end-to-end asr models,” in Proc. IEEE
- International Conference on Acoustics, Speech, and Signal Processing
- (ICASSP), 2020.
-
- |
-
-
-
-
-|
-[33]
- |
-
-D. Park, Y. Zhang, C.-C. Chiu, Y. Chen, B. Li, W. Chan, Q. Le, and Y. Wu,
- “Specaugment on large scale datasets,” in Proc. IEEE International
- Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2020.
-[ pdf ]
-
- |
-
-
-
-
-|
-[34]
- |
-
-T. Sainath, Y. He, B. Li, A. Narayanan, R. Pang, A. Bruguier, S. yiin Chang,
- W. Li, R. Alvarez, Z. Chen, C. cheng Chiu, D. Garcia, A. Gruenstein, K. Hu,
- M. Jin, A. Kannan, Q. Liang, I. McGraw, C. Peyser, R. Prabhavalkar,
- G. Pundak, D. Rybach, Y. Shangguan, Y. Sheth, T. Strohman, M. Visontai,
- Y. Wu, Y. Zhang, and D. Zhao, “A streaming on-device end-to-end model
- surpassing server-side conventional model quality and latency,” in
- Proc. IEEE International Conference on Acoustics, Speech, and Signal
- Processing (ICASSP), 2020.
-
- |
-
-
-
-
-|
-[35]
- |
-
-A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang,
- Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented transformer
- for speech recognition,” in Proc. Interspeech, 2020.
-[ pdf ]
-
- |
-
-
-
-
-|
-[36]
- |
-
-W. Han, Z. Zhang, Y. Zhang, J. Yu, C.-C. Chiu, J. Qin, A. Gulati, R. Pang, and
- Y. Wu, “Contextnet: Improving convolutional neural networks for automatic
- speech recognition with global context,” in Proc. Interspeech, 2020.
-[ pdf ]
-
- |
-
-
-
-
-|
-[37]
- |
-
-W. Li, J. Qin, C.-C. Chiu, R. Pang, and Y. He, “Parallel rescoring with
- transformer for streaming on-device speech recognition,” in Proc.
- Interspeech, 2020.
-
- |
-
-
-
-
-|
-[38]
- |
-
-D. S. Park, Y. Zhang, Y. Jia, W. Han, C.-C. Chiu, B. Li, Y. Wu, and Q. V. Le,
- “Improved noisy student training for automatic speech recognition,” in
- Proc. Interspeech, 2020.
-[ pdf ]
-
- |
-
-
-
-## Language understanding
-
-
-
-
-
-
-
-|
-[1]
- |
-
-J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen,
- Y. Zhang, Y. Wang, R. Skerry-Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and
- Y. Wu, “Natural TTS synthesis by conditioning WaveNet on mel spectrogram
- predictions,” in Proc. IEEE International Conference on Acoustics,
- Speech, and Signal Processing (ICASSP), 2018.
-[ sound examples |
-pdf ]
-
- |
-
-
-
-
-|
-[2]
- |
-
-J. Chorowski, R. J. Weiss, R. A. Saurous, and S. Bengio, “On using
- backpropagation for speech texture generation and voice conversion,” in
- Proc. IEEE International Conference on Acoustics, Speech, and Signal
- Processing (ICASSP), 2018.
-[ sound examples |
-pdf ]
-
- |
-
-
-
-
-|
-[3]
- |
-
-Y. Jia, Y. Zhang, R. J. Weiss, Q. Wang, J. Shen, F. Ren, Z. Chen, P. Nguyen,
- R. Pang, I. Lopez-Moreno, and Y. Wu, “Transfer learning from speaker
- verification to multispeaker text-to-speech synthesis,” in Advances in
- Neural Information Processing Systems, 2018.
-[ sound examples |
-pdf ]
-
- |
-
-
-
-
-|
-[4]
- |
-
-W. N. Hsu, Y. Zhang, R. J. Weiss, H. Zen, Y. Wu, Y. Wang, Y. Cao, Y. Jia,
- Z. Chen, J. Shen, P. Nguyen, and R. Pang, “Hierarchical generative modeling
- for controllable speech synthesis,” in Proc. International Conference
- on Learning Representations (ICLR), 2019.
-[ sound examples |
-pdf ]
-
- |
-
-
-
-
-|
-[5]
- |
-
-W. N. Hsu, Y. Zhang, R. J. Weiss, Y. A. Chung, Y. Wang, Y. Wu, and J. Glass,
- “Disentangling correlated speaker and noise for speech synthesis via data
- augmentation and adversarial factorization,” in NeurIPS 2018 Workshop
- on Interpretability and Robustness in Audio, Speech, and Language, 2018.
-[ pdf ]
-
- |
-
-
-
-
-|
-[6]
- |
-
-H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu,
- “LibriTTS: A corpus derived from LibriSpeech for text-to-speech,” in
- Proc. Interspeech, 2019.
-[ data |
-pdf ]
-
- |
-
-
-
-
-|
-[7]
- |
-
-F. Biadsy, R. J. Weiss, P. Moreno, D. Kanvesky, and Y. Jia, “Parrotron: An
- end-to-end speech-to-speech conversion model and its applications to
- hearing-impaired speech and speech separation,” in Proc. Interspeech,
- 2019.
-[ sound examples |
-pdf ]
-
- |
-
-
-
-
-|
-[8]
- |
-
-Y. Zhang, R. J. Weiss, H. Zen, Y. Wu, Z. Chen, R. J. Skerry-Ryan, Y. Jia,
- A. Rosenberg, and B. Ramabhadran, “Learning to speak fluently in a foreign
- language: Multilingual speech synthesis and cross-language voice cloning,”
- in Proc. Interspeech, 2019.
-[ sound examples |
-pdf ]
-
- |
-
-
-
-
-|
-[9]
- |
-
-G. Sun, Y. Zhang, R. J. Weiss, Y. Cao, H. Zen, and Y. Wu, “Fully-hierarchical
- fine-grained prosody modeling for interpretable speech synthesis,” in
- Proc. IEEE International Conference on Acoustics, Speech, and Signal
- Processing (ICASSP), 2020.
-[ sound examples |
-pdf ]
-
- |
-
-
-
-
-|
-[10]
- |
-
-G. Sun, Y. Zhang, R. J. Weiss, Y. Cao, H. Zen, A. Rosenberg, B. Ramabhadran,
- and Y. Wu, “Generating diverse and natural text-to-speech samples using a
- quantized fine-grained VAE and auto-regressive prosody prior,” in
- Proc. IEEE International Conference on Acoustics, Speech, and Signal
- Processing (ICASSP), 2020.
-[ sound examples |
-pdf ]
-
- |
-
-
-
-## Speech translation
-
-
-
-
-