unbabel-comet

High-quality Machine Translation Evaluation

These details have not been verified by PyPI

Project links

Download

GitHub Statistics

View statistics for this project via Libraries.io, or by using our public dataset on Google BigQuery

Project description

Note: This is a Pre-Release Version. We are currently working on results for the WMT2020 shared task and will likely update the repository in the beginning of October (after the shared task results).

Quick Installation

We recommend python 3.6 to run COMET.

Detailed usage examples and instructions can be found in the Full Documentation.

Simple installation from PyPI

pip install unbabel-comet

To develop locally:

git clone https://github.com/Unbabel/COMET
pip install -r requirements.txt
pip install -e .

Scoring MT outputs:

Via Bash:

Example:

echo -e "Hello world\nThis is a sample" >> src.en
echo -e "Oi mundo\neste é um exemplo" >> hyp.pt
echo -e "Olá mundo\nisto é um exemplo" >> ref.pt

comet score -s src.en -h hyp.pt -r ref.pt

You can export your results to a JSON file using the --to_json flag and select another model/metric with --model.

comet score -s src.en -h hyp.pt -r ref.pt --model wmt-large-hter-estimator --to_json segments.json

Via Python:

from comet.models import download_model
model = download_model("wmt-large-da-estimator-1719", "path/where/to/save/models/")
data = [
    {
        "src": "Hello world!",
        "mt": "Oi mundo!",
        "ref": "Olá mundo!"
    },
    {
        "src": "This is a sample",
        "mt": "este é um exemplo",
        "ref": "isto é um exemplo!"
    }
]
model.predict(data)

Simple Pythonic way to convert list or segments to model inputs:

source = ["Hello world!", "This is a sample"]
hypothesis = ["Oi mundo!", "este é um exemplo"]
reference = ["Olá mundo!", "isto é um exemplo!"]

data = {"src": source, "mt": hypothesis, "ref": reference}
data = [dict(zip(data, t)) for t in zip(*data.values())]

model.predict(data)

Model Zoo:

Model	Description
↑`wmt-large-da-estimator-1719`	RECOMMENDED: Estimator model build on top of XLM-R (large) trained on DA from WMT17, WMT18 and WMT19
↑`wmt-base-da-estimator-1719`	Estimator model build on top of XLM-R (base) trained on DA from WMT17, WMT18 and WMT19
↓`wmt-large-hter-estimator`	Estimator model build on top of XLM-R (large) trained to regress on HTER.
↓`wmt-base-hter-estimator`	Estimator model build on top of XLM-R (base) trained to regress on HTER.
↑`emnlp-base-da-ranker`	Translation ranking model that uses XLM-R to encode sentences. This model was trained with WMT17 and WMT18 Direct Assessments Relative Ranks (DARR).

QE-as-a-metric:

Model	Description
`wmt-large-qe-estimator-1719`	Quality Estimator model build on top of XLM-R (large) trained on DA from WMT17, WMT18 and WMT19.

Train your own Metric:

Instead of using pretrained models your can train your own model with the following command:

comet train -f {config_file_path}.yaml

Supported encoders:

Tensorboard:

Launch tensorboard with:

tensorboard --logdir="experiments/lightning_logs/"

Download Command:

To download public available corpora to train your new models you can use the download command. For example to download the APEQUEST HTER corpus just run the following command:

comet download -d apequest --saving_path data/

unittest:

pip install coverage

In order to run the toolkit tests you must run the following command:

coverage run --source=comet -m unittest discover
coverage report -m

Project details

These details have not been verified by PyPI

Project links

Download

GitHub Statistics

View statistics for this project via Libraries.io, or by using our public dataset on Google BigQuery

Release history Release notifications | RSS feed

2.2.2

Mar 13, 2024

2.2.1

Jan 8, 2024

2.2.0

Oct 23, 2023

2.1.1

Oct 13, 2023

2.1.0

Sep 21, 2023

2.0.2

Aug 3, 2023

2.0.1

Apr 5, 2023

2.0.0

Mar 13, 2023

1.1.3

Oct 4, 2022

1.1.2

Jun 6, 2022

1.1.1

Jun 1, 2022

1.1.0

Apr 2, 2022

1.0.1

Nov 19, 2021

1.0.0

Nov 19, 2021

1.0.0rc9 pre-release

Oct 21, 2021

1.0.0rc8 pre-release

Oct 18, 2021

1.0.0rc7 pre-release

Oct 18, 2021

1.0.0rc6 pre-release

Sep 28, 2021

1.0.0rc5 pre-release

Sep 4, 2021

1.0.0rc4 pre-release

Aug 16, 2021

1.0.0rc3 pre-release

Aug 15, 2021

1.0.0rc2 pre-release

Aug 10, 2021

1.0.0rc1 pre-release

Jul 27, 2021

0.1.0

Mar 11, 2021

0.0.7

Feb 9, 2021

0.0.6.post2

Nov 25, 2020

0.0.6.post1

Nov 24, 2020

0.0.6

Nov 21, 2020

This version

0.0.4

Oct 8, 2020

0.0.3

Sep 22, 2020

0.0.2

Sep 22, 2020

0.0.1 yanked

Sep 22, 2020

Reason this release was yanked:

missing MANIFEST with reqs

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

unbabel-comet-0.0.4.tar.gz (42.2 kB view hashes)

Uploaded Oct 8, 2020 Source

Hashes for unbabel-comet-0.0.4.tar.gz

Hashes for unbabel-comet-0.0.4.tar.gz
Algorithm	Hash digest
SHA256	`85e4ca54ee57bf61b329c51813ef82b2c8936ff550fd10b9ee44835c005e26b9`
MD5	`44693e3aef563d4fb73937463dd3c073`
BLAKE2b-256	`f5a890d99235f77a56df0a6ed8024e45089fd8c327a27fa4e345ffaa21250fed`