Finding the Active Voice

Last week, Harper hit a stroke of luck. It was fea­tured on MakeUseOf. The down­stream so­cial me­dia posts col­lec­tively gar­nered nearly 300,000 views and drove a sig­nif­i­cant amount of traf­fic to our site. This all hap­pened in the wake of LanguageTool shut­ting down the free edi­tion of their soft­ware. These two events com­pound to bring more at­ten­tion than ever to Harper — which is amaz­ing. It means our mis­sion and val­ues are res­onat­ing with peo­ple, which is al­ways a good thing.

Reading the MUO ar­ti­cle, we hear a lot of great things about Harper. We also hear that it skips many of the pre­mium bells and whis­tles of Grammarly”. The au­thor goes on to ex­plain that the pri­vacy Harper pro­vides more than makes up for any miss­ing pre­mium fea­tures, but the point is clear: those fea­tures are deeply de­sired.

So, the ques­tion be­comes: Which of Grammarly’s fea­tures should we work on first, and how?

The Active Voice

I pro­pose that we should fo­cus first on help­ing our users find the ac­tive voice. For con­text, the ac­tive voice is the style of writ­ing where sub­ject of the verb in a clause is the doer of the ac­tion. This is in con­trast to the pas­sive voice, which is where the sub­ject is the re­ceiver of the ac­tion. For ex­am­ple: the sen­tence, The postal car­rier was bit­ten by the dog” is writ­ten in the pas­sive voice, while the equiv­a­lent sen­tence, The dog bit the postal car­rier.” is writ­ten in the ac­tive voice.

Text writ­ten in the ac­tive voice is com­monly viewed to be more au­thor­i­ta­tive, con­fi­dent, and eas­ier to un­der­stand. Be­ing able to help users use the ac­tive voice is one of the most com­monly re­quested fea­tures in Harper, and in­clud­ing the fea­ture would be a huge step to­wards com­pet­ing di­rectly with Grammarly Premium.

How should we go about help­ing our users with their ac­tive voice?

How It Could Be Done

I spoke briefly with Matt, and we agreed that a two tier so­lu­tion would be best. A fast al­go­rithm or model would de­tect in­stances of the pas­sive voice, let­ting a larger more com­pu­ta­tion­ally ex­pen­sive model gen­er­ate a mod­i­fi­ca­tion in the ac­tive voice.

Fortunately, there is al­ready ex­ten­sive lit­er­a­ture on the de­tec­tion of the pas­sive voice. In par­tic­u­lar, I found the PassivePy pa­per stim­u­lat­ing. In fact, we can im­ple­ment their ideas quite eas­ily us­ing the Weir lan­guage al­ready baked into Harper. I have done so in a pri­vate branch. It turned out to be ~20 lines of code. That is pretty good bang for the buck!

The sec­ond piece, which has to do with the ac­tual sim­pli­fi­ca­tion of text and con­ver­sion from the pas­sive voice to the ac­tive voice is a tad more com­plex.

Matt and I agree that it will re­quire the use of a larger lan­guage model. The trou­ble is that it can­not be too large. Harper’s shtick is that it is fast, pri­vate, and that every­thing runs di­rectly on our user’s de­vices. That means whichever model we use for our style trans­fer will need to be rel­a­tively small.

I be­lieve the best so­lu­tion to this prob­lem is to take an off-the-shelf model, like one of Google’s T5 mod­els, and fine tune it for the spe­cific types of style trans­fer we need. These are rel­a­tively small mod­els (quantized, they can fit into spaces un­der 65 megabytes) and they run quite quickly, even on older hard­ware that does­n’t have ac­cess to ma­trix mul­ti­pli­ca­tion ac­cel­er­a­tors. There is prior art for run­ning this at 50 tok/​s in Chrome with­out WebGPU on a sin­gle core. The best part is that they’re un­der the Apache-2.0 li­cense!

How This Fits in with the Weirpack Project

These mod­els are small, but they’re not quite small enough to be a part of the stan­dard dis­tri­b­u­tion of Harper. I be­lieve this should be an opt-in fea­ture, and the best way to do that is to ex­pose the func­tion­al­ity via a Weirpack. If you don’t know what a Weirpack is, I highly sug­gest you read my pre­vi­ous blog posts on the sub­ject.

Everyone who wants this ad­di­tional func­tion­al­ity could just en­able it in the mar­ket­place. This con­tin­ues our goal to make Harper as cus­tomiz­able as our users want, while pro­vid­ing sen­si­ble de­faults.

What’s Next?

Once we have the sys­tem in place to de­tect and pro­vide sug­ges­tions for the ac­tive voice, we will be pre­pared to do other kinds of trans­for­ma­tion, like for ad­just­ing for­mal­ity or tone.

I’m re­ally ex­cited about this pro­ject, and I can’t wait to get started.

Published January 26, 2026 at 7:00 AM

Proofread by Harper.

Comments