ፊደል
ፊደል
ሀ ሉ ሒ ማ ሜ ር ሶ ሿ ቁ ቢ ታ ቼ ኅ ኖ ኚ ኣ ኬ
Currently Under Development
Full Release in 23 days (23d 00h 00m 00s)
02 / PERFORMANCE BENCHMARKS

Performance Lab

Interactive compiler benchmarks comparing raw execution loops across WebAssembly (compiled from Rust), standard JavaScript runtimes, and native Python C-extensions.

Rust C-Corewasm32-unknownpyo3 bindings

Product Quality Targets & Gaps

Current actual performance vs project shippable milestones

FeatureMetricActualShippableCompetitiveWorld-ClassStatus / Gap
NormalizerHomophone recall100.0%95%98%99.5%Exceeded
Sentence TokenizerF1 on boundaries68.8%90%95%98%Below Minimum (-21.2%)
Light StemmerAccuracy (correct root)32.1%70%82%90%+Below Minimum (-37.9%)
Stopword RemovalPrecision (no corruption)100.0%97%99%99.8%Exceeded
TransliteratorRound-trip accuracy100.0%92%97%99%Exceeded
Word TokenizerToken F1-85%92%96%Roadmap
POS TaggerAccuracy-85%91%95%Roadmap
NERF1 per entity type-75%85%92%Roadmap
SentimentMacro F1-72%82%88%Roadmap
API Latencyp95 response time< 1.0ms< 200ms< 100ms< 50msExceeded
API UptimeMonthly uptime99.99%99.5%99.9%99.95%Exceeded

Live Benchmark Sandbox

WASM ACTIVE
Quick Integration Packages
PNPM NPM PACKAGE
pnpm add @fidel-tools/core
PIP PYTHON PACKAGE
pip install fidel-tools

Performance Visual Graph

Comparative engine metrics profile

Payload Size

Medium Paragraphs (~200 chars)

የገንዘብ ሚኒስቴር ምክር ቤተ ከሃያ ዓመታት በፊት ያወጣውን የ ተጨማሪ እሴት ታክስ ቫት አዋጅን የሚተካ ረቂቅ ተዘጋጀ።
69,805 ops/s
JS
93,817 ops/s
WASM
149,115 ops/s
PY
100,000 ops/s
Goal

Pipeline Execution Flow Chart

Sequential stages processing Amharic Unicode streams

STAGE 01

Input Stream

Accepts raw Amharic Unicode strings. Sanitizes input boundary buffers.

STAGE 02

Normalize

Collapses homophones, normalizes labialized orthographies, and resolves character duplication.

STAGE 03

Tokenize

Identifies sentence boundaries and word-level token configurations with fallback logic.

STAGE 04

Stemmer

Applies affix stripping and context rules to extract morphologically clean semantic roots.

Gold-Standard Test Corpus

Accuracy metrics evaluated against 2,000 reference sentences

Our automated validation suite runs on every commit against independent, hand-labeled base cases across multiple linguistic categories to measure real-world performance.

Normalization (Homophones)
100.00%
Goal Target: 98.00%
Homophone recall (Exact Match: 67.05%)
Stemming (Correct Root)
32.10%
Goal Target: 82.00%
Light affix-removal match
Tokenization (Boundaries)
68.77%
Goal Target: 95.00%
Sentence split boundary F1

Linguistic Category Breakdowns

Normalization
Homophones100.0%
Labialization100.0%
Gemination5.3%
Clean Text81.4%
Stemming
Regular Affixes25.5%
Irregular Words77.6%
Ambiguous Roots15.2%
Tokenization (F1)
Standard Ends100.0%
Word Separators (፡)17.7%
Abbreviations100.0%

Linguistic Analysis & Failure Modes

Documented limitations and architectural trade-offs

1. Normalization (Gemination Collapsing Threshold)

The normalizer is configured with a gemination threshold of 2. When a character repeats 3+ times (e.g. ምምም), it collapses to 2 characters (ምም). However, the ground truth is completely un-geminated (having only 1 character, e.g. ). Because the normalizer only collapses down to the threshold (2) instead of fully de-geminating to 1 character, it fails the exact match comparison against the un-geminated ground truth. This is the expected, correct behavior of the threshold but explains the lower score.

2. Stemming (Ambiguous Roots & Morphotactics)

As a light stemmer using longest-match affix-removal, the engine lacks a complete morphological analyzer or root lexicon. It fails on ambiguous roots (e.g., stripping the leading በ- from በላ resulting in , or the leading ከ- from ከፈለ resulting in ፈለ) and morphotactic changes (e.g., vowel elision/epenthesis like ደብዳቤ inflecting and stemming to ደብድአብ).

3. Tokenization (Hulet Neteb ፡ as Sentence Boundary)

The language pack specifies the Amharic word separator (hulet neteb ) as a sentence boundary. In modern standard writing, separates words rather than sentences. Because the tokenizer splits sentences on every , paragraphs using hulet net med are over-segmented into word-level fragments, resulting in 0% exact match sentence accuracy.

Raw Performance Registry

Throughput and latency percentiles collected dynamically on hardware (v0.1.7)

Intel Core i5-1155G7 (8 cores)Node.js v22
Scale / ImplementationThroughputp50 Latencyp95 Latencyp99 LatencySpeedup
Short Payload (JS)581,156 ops/s1.26 μs2.39 μs3.47 μs0.79x
Short Payload (WASM)458,261 ops/s1.62 μs3.01 μs4.34 μs
Medium Payload (JS)69,805 ops/s13.30 μs17.27 μs25.24 μs1.34x
Medium Payload (WASM)93,817 ops/s10.29 μs12.16 μs13.74 μs
Large Payload (JS)7,358 ops/s130.85 μs148.82 μs264.87 μs1.30x
Large Payload (WASM)9,529 ops/s103.19 μs114.05 μs125.68 μs

Boundary-Crossing Overhead Gate

WebAssembly binaries run at near-native compile speeds. However, passing data between JS and WASM requires allocating heap memory and encoding/decoding strings to UTF-8 bytes. On short strings, this boundary-crossing overhead dominates the computation time. For medium-length paragraphs (~150 chars), Rust loops outpace JS, yielding a 35% speedup.