Cheap and reliable Node.js hosting starts at $3/month, and $1/month static HTML hosting

My implementation of the original transformer model (Vaswani et al.). I've additionally included the playground.py file for visualizing otherwise seemingly hard concepts. Currently included IWSLT pretrained models.

Stars: ✭ 411 (-34.55%)

Mutual labels: attention-is-all-you-need

Awesome Fast Attention

list of efficient attention modules

Stars: ✭ 627 (-0.16%)

Mutual labels: attention-is-all-you-need

Speech Transformer

A PyTorch implementation of Speech Transformer, an End-to-End ASR with Transformer network on Mandarin Chinese.

Stars: ✭ 565 (-10.03%)

Mutual labels: attention-is-all-you-need

Nmt Keras

Neural Machine Translation with Keras

Stars: ✭ 501 (-20.22%)

Mutual labels: attention-is-all-you-need

View All Similar Projects ➔

The Transformer model in Attention is all you need：a Keras implementation.

A Keras+TensorFlow Implementation of the Transformer: "Attention is All You Need" (Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin, arxiv, 2017)

Usage

Please refer to en2de_main.py and pinyin_main.py

en2de_main.py

This task is same as in jadore801120/attention-is-all-you-need-pytorch: WMT'16 Multimodal Translation: Multi30k (de-en) (http://www.statmt.org/wmt16/multimodal-task.html). We borrowed the data proprocessing step 0 and 1 in the repository, and then construct the input file en2de.s2s.txt

Results

The code achieves near results as in the repository: about 70% valid accuracy. If using smaller model parameters, such as layers=2 and d_model=256, the valid accuracy is better since the task is quite small.

For your own data

Just preproess your source and target sequences as the format in en2de.s2s.txt and pinyin.corpus.examples.txt.

Some notes

For larger number of layers, the special learning rate scheduler reported in the papar is necessary.
In pinyin_main.py, I tried another method to train the deep network. I train the first layer and the embedding layer first, then train a 2-layers model, and then train a 3-layers, etc. It works in this task.

Upgrades

Reconstruct some classes.
It is more easier to use the components in other models, just import transformer.py
A fast step-by-step decoder is added, including an upgraded beam-search. But they should be modified to be reuseable.

Acknowledgement

Some model structures and some scripts are borrowed from jadore801120/attention-is-all-you-need-pytorch.

Note that the project description data, including the texts, logos, images, and/or trademarks, for each open source project belongs to its rightful owner. If you wish to add or remove any projects, please contact us at [email protected].

Stars: ✭ 628

Visit Git Page 🔗Visit User Page 🔗Visit Issues Page (24) 🔗