Kyubyong / Transformer

Licence: apache-2.0

A TensorFlow Implementation of the Transformer: Attention Is All You Need

Programming Languages

python

139335 projects - #7 most used programming language

perl

6916 projects

shell

77523 projects

Projects that are alternatives of or similar to Transformer

Sockeye

Sequence-to-sequence framework with a focus on Neural Machine Translation based on Apache MXNet

Stars: ✭ 990 (-72.85%)

Mutual labels: translation, attention-mechanism, attention-is-all-you-need, transformer

Pytorch Original Transformer

My implementation of the original transformer model (Vaswani et al.). I've additionally included the playground.py file for visualizing otherwise seemingly hard concepts. Currently included IWSLT pretrained models.

Stars: ✭ 411 (-88.73%)

Mutual labels: attention-mechanism, attention-is-all-you-need, transformer

Pytorch Transformer

pytorch implementation of Attention is all you need

Stars: ✭ 199 (-94.54%)

Mutual labels: translation, attention-is-all-you-need, transformer

Nmt Keras

Neural Machine Translation with Keras

Stars: ✭ 501 (-86.26%)

Mutual labels: attention-mechanism, attention-is-all-you-need, transformer

pynmt

a simple and complete pytorch implementation of neural machine translation system

Stars: ✭ 13 (-99.64%)

Mutual labels: translation, transformer, attention-mechanism

visualization

a collection of visualization function

Stars: ✭ 189 (-94.82%)

Mutual labels: transformer, attention-mechanism

OverlapPredator

[CVPR 2021, Oral] PREDATOR: Registration of 3D Point Clouds with Low Overlap.

Stars: ✭ 293 (-91.96%)

Mutual labels: transformer, attention-mechanism

FragmentVC

Any-to-any voice conversion by end-to-end extracting and fusing fine-grained voice fragments with attention

Stars: ✭ 134 (-96.32%)

Mutual labels: transformer, attention-mechanism

Transformer-in-Transformer

An Implementation of Transformer in Transformer in TensorFlow for image classification, attention inside local patches

Stars: ✭ 40 (-98.9%)

Mutual labels: transformer, attention-mechanism

NLP-paper

🎨 🎨NLP 自然语言处理教程 🎨🎨 https://dataxujing.github.io/NLP-paper/

Stars: ✭ 23 (-99.37%)

Mutual labels: transformer, attention-mechanism

Image-Caption

Using LSTM or Transformer to solve Image Captioning in Pytorch

Stars: ✭ 36 (-99.01%)

Mutual labels: transformer, attention-mechanism

Transformer Tensorflow

TensorFlow implementation of 'Attention Is All You Need (2017. 6)'

Stars: ✭ 319 (-91.25%)

Mutual labels: translation, transformer

speech-transformer

Transformer implementation speciaized in speech recognition tasks using Pytorch.

Stars: ✭ 40 (-98.9%)

Mutual labels: transformer, attention-is-all-you-need

transformer

A PyTorch Implementation of "Attention Is All You Need"

Stars: ✭ 28 (-99.23%)

Mutual labels: transformer, attention-is-all-you-need

enformer-pytorch

Implementation of Enformer, Deepmind's attention network for predicting gene expression, in Pytorch

Stars: ✭ 146 (-96%)

Mutual labels: transformer, attention-mechanism

dodrio

Exploring attention weights in transformer-based models with linguistic knowledge.

Stars: ✭ 233 (-93.61%)

Mutual labels: transformer, attention-mechanism

transformer

Neutron: A pytorch based implementation of Transformer and its variants.

Stars: ✭ 60 (-98.35%)

Mutual labels: transformer, attention-is-all-you-need

linformer

Implementation of Linformer for Pytorch

Stars: ✭ 119 (-96.74%)

Mutual labels: transformer, attention-mechanism

attention-is-all-you-need-paper

Implementation of Vaswani, Ashish, et al. "Attention is all you need." Advances in neural information processing systems. 2017.

Stars: ✭ 97 (-97.34%)

Mutual labels: transformer, attention-is-all-you-need

galerkin-transformer

[NeurIPS 2021] Galerkin Transformer: a linear attention without softmax

Stars: ✭ 111 (-96.96%)

Mutual labels: transformer, attention-mechanism

View All Similar Projects ➔

[UPDATED] A TensorFlow Implementation of Attention Is All You Need

When I opened this repository in 2017, there was no official code yet. I tried to implement the paper as I understood, but to no surprise it had several bugs. I realized them mostly thanks to people who issued here, so I'm very grateful to all of them. Though there is the official implementation as well as several other unofficial github repos, I decided to update my own one. This update focuses on:

readable / understandable code writing
modularization (but not too much)
revising known bugs. (masking, positional encoding, ...)
updating to TF1.12. (tf.data, ...)
adding some missing components (bpe, shared weight matrix, ...)
including useful comments in the code.

I still stick to IWSLT 2016 de-en. I guess if you'd like to test on a big data such as WMT, you would rely on the official implementation. After all, it's pleasant to check quickly if your model works. The initial code for TF1.2 is moved to the tf1.2_lecacy folder for the record.

Requirements

python==3.x (Let's move on to python 3 if you still use python 2)
tensorflow==1.12.0
numpy>=1.15.4
sentencepiece==0.1.8
tqdm>=4.28.1

Training

STEP 1. Run the command below to download IWSLT 2016 German–English parallel corpus.

bash download.sh

It should be extracted to iwslt2016/de-en folder automatically.

STEP 2. Run the command below to create preprocessed train/eval/test data.

python prepro.py

If you want to change the vocabulary size (default:32000), do this.

python prepro.py --vocab_size 8000

It should create two folders iwslt2016/prepro and iwslt2016/segmented.

STEP 3. Run the following command.

python train.py

Check hparams.py to see which parameters are possible. For example,

python train.py --logdir myLog --batch_size 256 --dropout_rate 0.5

STEP 3. Or download the pretrained models.

wget https://dl.dropbox.com/s/4lom1czy5xfzr4q/log.zip; unzip log.zip; rm log.zip

Training Loss Curve

Learning rate

Bleu score on devset

Inference (=test)

python test.py --ckpt log/1/iwslt2016_E19L2.64-29146 (OR yourCkptFile OR yourCkptFileDirectory)

Results

Typically, machine translation is evaluated with Bleu score.
All evaluation results are available in eval/1 and test/1.

tst2013 (dev)	tst2014 (test)
28.06	23.88

Notes

Beam decoding will be added soon.
I'm going to update the code when TF2.0 comes out if possible.

Note that the project description data, including the texts, logos, images, and/or trademarks, for each open source project belongs to its rightful owner. If you wish to add or remove any projects, please contact us at [email protected].

Cheap and reliable Node.js hosting starts at $3/month, and $1/month static HTML hosting

Kyubyong / Transformer

Programming Languages

Labels

Projects that are alternatives of or similar to Transformer

[UPDATED] A TensorFlow Implementation of Attention Is All You Need

Requirements

Training

Training Loss Curve

Learning rate

Bleu score on devset

Inference (=test)

Results

Notes