The post reads something like: looking for people for small paid online tasks, data entry, blog posting, etc. An honest description of a real need. You have 1,400 rows of exported CRM junk to normalise, 60 old posts to load into WordPress, a folder of supplier PDFs whose totals need typing into a sheet. None of it justifies a job posting, and none of it is worth your Saturday.
So you go looking for cheap hands. Either nobody replies who you would trust with a spreadsheet, or twelve people reply and you have no way to choose, or somebody delivers 1,400 completed rows in nine hours and about 300 are wrong in ways you discover in March.
Piecework outsourcing means paying per completed unit of work rather than per hour, and it fails on quality control far more often than it fails on labour supply. Pay per task without designing verification and you get exactly the quality you paid to measure, which is none. The fix is four things borrowed from data annotation practice: define the unit, write a spec a stranger cannot misread, seed a set of records whose answers you already know, and accept or reject the batch by random sample instead of by trust.
First decide whether the task is piecework at all
Three questions. A no on any one means you are buying judgment, and judgment bought by the unit from a stranger is the expensive kind.
Is the output objectively checkable? Could a second person holding only your spec, knowing nothing about your business, independently produce the same answer? “Type the invoice total into column F” passes. “Summarise what this customer wants” does not, because two careful people produce two different summaries and neither is wrong.
Is the unit small and uniform? One row, one record, one minute of audio, the same shape 500 times. If unit 300 needs a different approach from unit 12, you have a project, and projects need someone who accumulates context.
Does it need context about your business? If the right answer depends on knowing Henderson is on legacy pricing, the task is not separable from you, and every question the worker must ask is a unit of the attention you were buying back.
Then the tasks that pass a casual glance and belong nowhere near a one-off worker. Judgment about your customers: review responses, support replies, outreach in your name. Bookkeeping categorisation, which looks like data entry but is judgment in a spreadsheet costume. Anything needing your admin login or a sheet of named customers, because the access decision inside it is bigger than the task. And anything publishing publicly with no review gate: loading 60 posts into WordPress is piecework, loading and publishing them is not. The unit ends at saved as draft, correct category, images attached, alt text present. You press publish.
Where a task straddles the line, split it until the judgment sits on your side. “Build me a list of good-fit prospects” is not outsourceable. “For each of these 200 named companies, find the current head of finance on LinkedIn or the company site, record name, exact title and source URL, leave email blank” is. You kept the fit decision and gave away the typing. Do that before deciding where to post the job.
Where to outsource small online tasks, honestly compared
Micro-task platforms. Mechanical Turk, Clickworker and similar are built for this shape: atomised, verifiable, high volume, no context. Budget the fees. Amazon’s published requester pricing is 20% of the reward, an additional 20% on tasks with ten or more assignments, plus 5% for the Masters qualification. Good for image tagging, classification and short extraction from public sources, useless for anything needing your CMS.
Freelance marketplaces. Upwork, Fiverr, PeoplePerHour. One accountable human, escrow, a readable history, a dispute process that exists. The economics break at the bottom, where briefing and screening four applicants costs more than a $40 job is worth, so batch several small jobs into one contract.
Task-specific vendors. Transcription services, data-entry shops, image annotation providers, priced per minute, per record or per image. The good ones already run the quality layer described below, so you are buying spec discipline rather than hours.
A recurring part-time person, or a BPO team. Neither is piecework, and both become the right answer sooner than buyers admit. By the fifth time you brief a stranger on a monthly batch you have spent more than a part-timer would have cost, and at volume a vendor sells the one thing a one-off worker cannot: a contracted accuracy level.
Community and beer-money forums. Where the original question was asked, so answer it straight: the highest-variance supply you will find. Capable people, students and bots, no escrow, no identity verification, no recourse, and someone has to go first on a sum too small for either side to chase. Fine for zero-risk, fully checkable, tiny batches. Send nothing you would mind losing.
The spec is the highest-leverage thing you will write
An hour on the spec saves more rework than any amount of screening. Most small buyers write four lines in a chat message, then read the result as a comment on the worker’s diligence. Annotation teams write guidelines documents for the opposite reason: ambiguity in the instruction becomes disagreement in the output, and disagreement you did not anticipate is indistinguishable from carelessness. Seven sections, a page or two.
- The unit of work. What counts as one, since it is also your pricing unit and your defect unit. “One company row, complete when all five fields are filled or marked unresolvable.”
- One worked correct example. A real record, every field filled, as it should appear in the output file. Shown, not described.
- Three worked wrong examples, with reasons. The section nobody writes, and it does more work than the rest combined. For list building: title recorded as “Finance Director” when the site says “Group Finance Director” (copy titles exactly, never normalise); email recorded as j.smith@acme.com with no source (guessed from a pattern, the worst failure on this task, leave blank instead); source URL recorded as acme.com (needs the exact page, not the domain). Wrong examples define the boundary. Correct examples only define the centre.
- Edge-case decision rules. An if-then list from cases you know exist. Two people share the title. The site redirects because the company was acquired. The PDF total includes tax and your column does not.
- Source of truth when fields conflict. “Company website over LinkedIn over directory listing. Where website and LinkedIn disagree on title, record the website version and flag the row.”
- What to do when the record is ambiguous. A written default, in the imperative, because its absence is what produces invented data: do not guess. Leave the field blank, put “unclear” in the flag column, add one line in notes saying what was unclear, move on. Then honour it. Complain about blanks once and you have taught the worker to guess.
- Done means, and the delivery format. Exact columns in exact order, date format, file type, destination. Half the rework on data tasks is format.
Seed a gold standard, then judge the batch by sample
Before the batch goes out, do 20 to 40 units yourself, carefully, and keep the answers. Scatter them through the work, indistinguishable from the rest, saying nothing about which they are, and score those first when it comes back. You now have a measured accuracy figure rather than an impression. Two rules keep it honest. No clustering at the start and no suspiciously tidy gold records. And refresh the set between batches with a recurring worker, or they learn the set rather than the task.
Measuring beats trusting, and it works with genuinely non-expert labour. Snow, O’Connor, Jurafsky and Ng’s 2008 EMNLP paper scored paid non-expert annotations against expert gold labels on five language tasks and found high agreement on all five, needing roughly four non-expert annotations per item to emulate one expert on affect recognition. Take the method rather than the numbers: known answers, measured agreement, aggregation where agreement is weak. Sambasivan and colleagues at Google priced the alternative in their 2021 CHI study of 53 AI practitioners, reporting compounding downstream failures from upstream data problems at 92% prevalence and calling them invisible, delayed and often avoidable. Your CRM is lower stakes than cancer screening. The failure has the same shape.
Then sample the batch, randomly. Checking the first ten rows is worse than useless, because care is front-loaded on nearly every piecework task: people start slow and careful and speed up. The formal method is acceptance sampling by attributes, standardised as ISO 2859-1, Sampling procedures for inspection by attributes, Part 1: Sampling schemes indexed by acceptance quality limit (AQL) for lot-by-lot inspection. Declare a lot and an acceptance quality limit, read off a sample size and an acceptance number, inspect that many units, and reject the whole lot above the acceptance number. The spreadsheet version:
- Define the defect before you look. Per row or per field, stated in the spec. Is one wrong character in a postcode a defective row? Decide now, not while arguing over an invoice.
- Fix the sample size and reject number up front, in writing. For 500 rows: 50 random rows, every field checked, more than two defective rows means rework at no extra cost. Writing it down in advance is what makes it a standard rather than a mood.
- Reject the lot, not the rows. Returning only the errors you found leaves the ones you missed, and whole-batch rework is the incentive that makes the first pass careful.
- If it passes, stop checking. A pass you ignore means you bought inspection, not work.
Where an error is genuinely expensive, add a second independent pass on those fields only and reconcile the disagreements. Clinical data management has used double entry for decades on the same reasoning: it doubles cost on the fields you apply it to and catches what sampling structurally misses, the quiet, plausible, evenly spread mistake. Use it on the column that will receive 200 emails, and single-enter the rest.
Per-task pricing is a quality setting, not just a price
Pay per unit with no verification and you have not paid for correct units. You paid for units that look complete, because completeness is the only property measured. A rational worker fills every cell fast, guesses plausibly where the source is hard, and flags nothing, since flagging looks like failure.
The signature failure in list building is fabricated contact data, and it is rarely malice. Someone cannot find an email, no rule tells them what to do, and they know the pattern is usually first initial plus surname at the domain. So they type it, and the row looks perfect. Three tells: the pattern is uniform down the column, generic catch-all addresses appear wherever the real answer was hard, and the source URL column is empty or points at a homepage. Two defences. Require a source URL per researched field, because a fabricated value has nowhere to point. And run the emails through a verification service, where an unusual bounce rate on one batch is your answer.
Per-hour pricing with no verification is not the fix either, since it buys elapsed time and sets no accuracy target. Use three things together: a per-unit rate, a written accuracy threshold carrying a rework obligation, and a completion bonus that pays when the gold set passes. Announce the threshold before work starts, because a standard produced afterwards reads as a way to avoid paying.
Which brings up rates, where the case for paying properly is mechanical rather than moral. Hara and colleagues analysed 2,676 workers completing 3.8 million tasks on Mechanical Turk for CHI 2018 and found a median hourly wage of roughly $2, with only 4% earning more than $7.25 an hour, while requesters paid an average of over $11, the gap going to unpaid time searching for work and on rejected or unsubmitted tasks. Read it as incentive design. Work priced so a careful pass earns two dollars an hour selects for speed over accuracy, because at that rate accuracy is unaffordable to the worker. You pay twice, in the reward and in the rework. A sensible rate is a quality control measure, and it gets you the same person next time.
Access, data, and getting people paid
- Never share your login. Create a separate account with the least privilege that completes the task. In WordPress, Contributor or Author, never Editor or Administrator.
- Share the file, not the folder. One spreadsheet, named people only, commenter access where the task allows, link sharing off.
- No customer personal data in a sheet sent over chat. Strip it, send record IDs and the fields the task needs, rejoin on your side. A task that cannot run without identifiable personal data needs processor terms, which is a different conversation from a $40 job.
- Revoke the same day. Delete the account rather than deactivating it, rotate anything shared, pull the file from shared drives. Four minutes, and nobody does it.
Paying small amounts across borders carries friction that surprises first-time buyers. A marketplace handles it inside escrow and takes a cut for doing so; paying directly means a transfer service where a six-dollar payout is absurd on fees. Batch to weekly payments, agree the method and who absorbs the fee before work starts, and never ask a worker to eat a charge worth a fifth of their payment. Get an invoice even for tiny amounts.
When piecework should stop being piecework
The signal is arithmetic. Multiply the hours you spend specifying, chasing, sampling and reworking by what an hour of your time is worth, then compare that to what you paid for the work. When the first number is bigger you bought a supervision job rather than capacity, and the arrangement has failed even if every batch eventually came out clean.
Three softer signals. You are rewriting the spec substantially every batch, meaning the task is unstable and needs a person who accumulates the knowledge rather than a document attempting to. Volume has reached thousands of units a month, where a vendor will contract an accuracy level and run the sampling themselves. Or the task grew a tentacle into customer contact or credentialed systems, in which case reread the first section, because it left the category. Nothing is wasted: the spec, gold set and sampling plan are what you hand the vendor, and they separate a proposal priced against your standard from one priced against a vague description.
If the work has outgrown one-off hands, that is what AB7 Solutions runs as a service: data annotation and labelling, AI data operations with human-in-the-loop review, and back-office and BPO teams for data entry, list building, transcription and CMS production, with gold-standard sets and sample-based QC inside the process rather than left to the buyer. Send a spec and a sample batch and we will say whether it needs a team or a better sampling plan. Call +1 321 341 7733, email ab@ab7solutions.com or director@ab7solutions.com, or see what the data and back-office teams cover at www.ab7solutions.com.
Questions that come up next
How big should the gold set be? Enough that one error does not swing the result: twenty to forty units, or 5% of the batch, whichever is larger. With ten, one mistake reads as 90% accuracy and two read as 80%, too coarse to decide with.
Should I tell the worker there is a gold set and a sampling plan? Yes, which does not contradict keeping the gold units unidentifiable. The known standard is the entire incentive: someone who knows accuracy is measured against known answers, and that the batch can be rejected whole, behaves differently from someone assuming nobody checks. Hide the units, publish the rule.
Can I use AI for these tasks instead? For some, better than piecework. Extraction from clean structured documents, classification against a fixed category set and first-pass transcription are all reasonable, and your spec plus gold set plus sampling plan is the test rig for judging whether the output is good enough. It does not remove the verification work. It changes who does the first pass.
Sources: Snow, O’Connor, Jurafsky & Ng, “Cheap and Fast, But is it Good?”, EMNLP 2008; Sambasivan et al., “Data Cascades in High-Stakes AI”, ACM CHI 2021; Hara et al., “A Data-Driven Analysis of Workers’ Earnings on Amazon Mechanical Turk”, ACM CHI 2018; Amazon Mechanical Turk requester pricing; ISO 2859-1 (1999 edition withdrawn January 2026, replaced by ISO 2859-1:2026).