What crowdsourcing gets right and wrong
Crowdsourcing Collecting information, work or contributions from many people instead of relying on one central source or team. Think of it like Instead of one person trying to map every pothole in a city, thousands of drivers can report the ones they encounter. Example Pumply lets drivers contribute fuel prices and station information they observe in the real world. sounds simple when you describe it from far away: let people contribute information, let enough people participate, and the product gets better as the community grows. I still believe in that optimistic version. A person standing at a fuel station can know something that no centralized database knows yet: the price changed this morning, the station is temporarily closed, the brand changed, or a place shown on an old map no longer exists. That kind of local knowledge is powerful. But building Pumply has made me much less comfortable with another idea that often gets smuggled into the word crowdsourcing: that enough contributions somehow become truth automatically. A crowd can give you coverage and freshness, but it can also give you duplicate records, stale assumptions, spam, coordinated mistakes and ten people repeating the same wrong thing. The interesting problem is therefore not how to collect more submissions; it is how to turn submissions into evidence without pretending every submission deserves the same amount of trust.
The crowd is useful because it is already there
Michael Goodchild described volunteered geographic information as a world in which ordinary people become something like distributed geographic sensors. That framing makes sense to me because it explains what crowdsourcing is good at. A centralized organization can publish a national dataset, but it cannot physically be everywhere at once. A community already is. For Pumply, this matters because fuel information changes in the real world before it changes in a database. A station may change its price, close for maintenance or rebrand while old map data still carries the previous identity. The person passing that station has an information advantage, and crowdsourcing lets the product capture it. This is the part crowdsourcing gets very right: it moves observation closer to the event.
An observation is not the same thing as truth
The danger begins when every contribution is stored as if it were a final fact. If someone reports that petrol is GH₵x at a station, what do I actually know? I know that a particular user or device submitted a particular value at a particular time for a particular station. That is already useful, but it is not identical to knowing that the price is definitely GH₵x. Maybe the person typed the diesel price into the petrol field, looked at an old sign, reported after the pump price changed, or was not at the station at all. Maybe the report is correct. The database should not erase those possibilities just because the interface wants one clean number.
This is why I increasingly think of crowdsourced data as a stream of claims with context, not a magical shared spreadsheet of facts. A claim has a source and a time; it can disagree with another claim, become stale or be corroborated Supported by another independent piece of evidence. Think of it like One person says it rained; a second witness and a wet road make the claim stronger. Example A driver's rebrand report becomes more convincing if another nearby report or an official station listing independently shows the same new brand. later. Sometimes the right system response is not “pick a winner immediately” but “we have conflicting evidence.” That sounds less elegant, but it is closer to reality.
Verifiability matters more than popularity
OpenStreetMap has a principle I keep returning to: data should, as far as reasonably possible, be verifiable Possible for someone else to check against observable evidence. Think of it like A claim like 'this shop closes at 8 p.m.' can be checked against the posted hours; a claim like 'this is the best shop' cannot be verified in the same way. Example Another mapper should be able to visit a station and check whether the brand shown in the data matches the sign at the forecourt. . Another mapper should be able to observe the same thing and determine whether the recorded information is true or false. I like that principle because it gives a crowdsourced system a standard that is stronger than “a lot of people said it.” Ten people can repeat a rumour; one person standing in front of a price board can make a directly verifiable observation. Those are not equivalent forms of evidence.
This is also why I do not think every crowdsourced product should be built around majority voting. Majority agreement is useful in some systems, but it can become misleading when contributors are copying the same source, one location has far more active users than another, or a mistake becomes popular before anyone checks the real world. For Pumply, I care more about where the report came from, when it happened and what it can be checked against than about pretending every problem can be solved with upvotes.
Trust rules should remove specific bad behaviours
Pumply does not let a person sit anywhere and update any station. For a price report, the current rule requires the reporter to be within 200 metres of the station, and the backend checks the distance again with PostGIS. That does not prove the report is correct; it proves something narrower and still useful: the reporter was plausibly close enough to have observed the station. I think good trust systems often work like this. They do not prove the whole truth at once; they remove specific classes of bad behaviour. A proximity rule makes remote mass-editing harder, a quota makes spam more expensive, a timestamp exposes age, an identity gives repeated behaviour continuity, and a review queue lets unusual additions wait for inspection. None of these mechanisms creates perfect data, but together they make bad data harder to create accidentally and harder to create at scale.
Freshness needs the same kind of specificity. If somebody reports that a station is open, that observation can become stale quickly. A petrol price may remain useful for much longer, and the physical existence of a station can remain valid for months or years. So a crowdsourcing system needs to ask not merely “is this old?” but what kind of fact is this, and how quickly can this kind of fact stop being true? That is why Pumply treats availability, price and station identity differently instead of assigning one global expiry time to everything.
Provenance is part of the value
Provenance The record of where data came from and what happened to it over time. Think of it like It is a chain of custody for information. Example A price can carry who reported it, when it was reported, whether an admin changed it later and which source the current value came from. is useful because the W3C’s PROV model is built around a simple idea: information about how data was produced helps people judge its quality, reliability and trustworthiness. Pumply does not implement the PROV standard directly, but the idea changed how I think about product data. A value without a source is weaker than the same value with a source; a value without a timestamp is weaker than the same value with a timestamp; a station identity with no explanation of how it changed is harder to trust than one with an audit trail showing that it was renamed or rebranded.
This becomes especially important when official and community data disagree. If a fuel brand publishes a national price and a driver reports a different price at a particular station, I do not want to destroy one piece of evidence because the other one exists. The brand price tells me what the company says; the station report tells me what somebody observed at that location. The useful product question is not always “which database row wins?” Sometimes it is “which source should the user see for this decision, and how clearly can I explain where it came from?”
The contributor should not have to understand the trust system
There is a product-design tension here. The backend can track provenance, rate limits, distance, report age, conflicts, moderation state and audit history, while the person contributing should still be able to do something as simple as: I am here. This is the price I see. Submit. If contributing feels like filing a government form, the community will not contribute enough data for the system to become useful. So the product has to absorb the complexity on the contributor’s behalf. The contributor does not need to know how PostGIS calculates distance or how the quota table works; they need a clear explanation of what the product expects and fair feedback when a rule blocks them.
This becomes more important as the community grows because growth does not only create more data, it creates more disagreement. Two people can report different prices minutes apart, two map sources can disagree about a station name, a community report can disagree with a brand’s official price, and a real forecourt can change brand while older datasets remain correct about the past. At small scale these look like edge cases. At larger scale, disagreement becomes a normal operating condition. That is why Pumply’s pending stations, corrections, activity, price history and audit records are not back-office decoration; they are the machinery that lets the product keep functioning when sources disagree.
Crowdsourcing should complement authority, not imitate it
I do not want Pumply to pretend the crowd is a regulator, an oil company or an official national registry. Those sources have different kinds of authority. The crowd has a different advantage: observation from the ground. The strongest system can use both. Official rules can explain how fuel pricing is structured, brand publications can provide a fallback or reference price, seed datasets can provide coverage, community reports can provide local freshness, and admin review can reconcile conflicts. The product becomes useful because the sources are allowed to remain different instead of being flattened into one fake source of truth.
This is the broader lesson I am taking from Pumply. Crowdsourcing is not the absence of infrastructure; in some ways it needs more infrastructure around trust because the inputs are decentralized Coming from many separate people or places rather than one single controlling source. Think of it like A radio station broadcasts from one centre; a neighbourhood WhatsApp group gets updates from many different phones. Example In crowdsourcing, reports arrive from many independent users instead of one national data-entry office. . The crowd gives you observations. The product still has to decide how those observations become information. That is the part I find most interesting now.
References
Michael F. Goodchild — Citizens as sensors: the world of volunteered geography ↗
OpenStreetMap Wiki — Verifiability ↗
W3C — PROV Model Primer ↗