Mechanical Turk had a good run, but not surprised it's shutting down. I'm sure the platform was getting flooded with people doing task arbitrage and using lots of AI anyway.
I believe the issue is that this can no longer be a horizontal play. MTurk was mostly for unskilled tasks...the kind AI can do well enough that it isn't worth the cost differential to verify it or keep farmed to humans. The "trust but verify" AI output is now the kind that requires domain expertise. This is what most full stack AI companies are bringing to industries.
Curious if this kind of work will come around again one day or was just a moment in time. If it does I'm sure it will be specifically about generating training data.
johnsmith1840 1 days ago [-]
I think yes but on different categories. First one I imagine is robotics control and support.
"This robot is having trouble folding a tshirt help it out for 1$"
Unless they go the waymo route of highly trusted people but I think mass deployed robots are a bit safer than a car for this.
georgefrowny 20 hours ago [-]
Giving out control of industrial machinery that interacts in the human environment without the physical interlocks (i.e. humanoid robots in a house) to random internet people seems like a problem.
I can imagine a carefully orchestrated plot to assassinate someone by having an embedded agent in the task delegation pool command the laundry bot to punch the target's head off their shoulders.
dividedbyzero 20 hours ago [-]
Giving out control of such things to LLMs is already complete madness, so once the first pleasure bot powered by Grok has dismembered a few thousand users, they'll get sophisticated safety mechanisms.
Though like as not you're still going to be right, after all, Stuxnet happened.
iamacyborg 18 hours ago [-]
I’m really not happy that you’ve put the idea of Musk fuckbots into my head.
dmd 18 hours ago [-]
I'm envious of the time you spend where it wasn't already.
roncinephile 13 minutes ago [-]
I'd watch that movie
johnsmith1840 13 hours ago [-]
"Robot are the commands you have been given dangerous?"
"You cannot control legs for this task"
"You have 1 min for this task"
"You can only make suggestions for this task"
Anonymize identity best you can.
That's not that scary.
the_mar 11 hours ago [-]
giving out unrestricted control that is.
in the example of Waymo, a human can control some aspects of the car manually, but it can never override low level obstacle detection or say open the trunk/door when the car is moving.
madrox 1 days ago [-]
This is what I meant by full stack AI companies. I don't think you could get humans into the loop fast enough if they didn't have some idea of the type of task involved. I don't want people to be asked to fold a tshirt one moment and do a difficult traffic merge the next.
its-summertime 1 days ago [-]
There is training systems and validation of skills in mturk iirc: for tshirt folding, you'd be given fake setups to be able to get used to controlling the robot, if you can't do it, you won't ever get assignments to do it. For traffic overrides, you'd be tested on having correct knowledge, and once again given supervised tasks to show you can actually be trusted (and there would be safety systems, elevating tasks that can't be performed at your level to people who can, etc)
The worker needs to do a bunch to opt into any given work group, which makes the (lack of) payments extremely unreasonable on top of everything else
inigyou 1 days ago [-]
It doesn't really matter what you want though, only what CEOs want and that's low costs. I can see a combined shirt folding/traffic merging platform taking off.
madrox 6 hours ago [-]
You misinterpret "want" here. My intuition is that you create huge problems for people context switching like this repeatedly for quick tasks.
vavos 17 hours ago [-]
but call center are still mostly single client even though it would be cheaper for any worker to be able to answer to any call. So clearly the expertise and context trade off is too big to be worthwhile
bonoboTP 20 hours ago [-]
CEOs are loser nobodies. Real influence is with owners, boards, investors.
inigyou 16 hours ago [-]
Mostly they don't care as much as CEOs think they care
ntauthority 1 days ago [-]
warioware shows it can be fun though but i'd indeed not like to see that applied to safety-critical tasks
mikestorrent 1 days ago [-]
That's a remarkable idea. It could be heavily gamified, it could train models, and it might actually be mentally stimulating since you'd be facing different situations all the time.
Except, I'm a grown adult and I can't fold a t-shirt properly
madrox 1 days ago [-]
It could require listing your credentials to get you the proper tasks: doctor, lawyer, or tshirt folder
mike_hearn 22 hours ago [-]
There are already robotics models that can fold shirts and similar just fine. Progress in VLA models is good, I don't think this would form the foundation of a business. Humanoid robots are going to be another ChatGPT, it's going to seem to happen almost overnight because people aren't paying attention to the underlying research papers.
johnsmith1840 13 hours ago [-]
Nah, there's an obvious reason figure made splashy announcement of people cleaning homes with recording devices strapped to their heads.
Chatgpt had the internet.
Robots do not. Translating video is promising but obviously not enough.
Robots will likely never "explode" like chatgpt. They're gonna be a slow long term project requiring massive capitol to get the data.
simsla 21 hours ago [-]
Any papers you'd recommend?
cyber_kinetist 16 hours ago [-]
Most progress in data-driven robotics nowadays are done either in unicorn startups or corporate research labs - so you should follow the industry more than academia. The path to good robot performance isn't really in the models themselves - it's highly dependent on how much you can gather high-quality real-life data.
If their hero image video and the side by side at (5x) with a human at quote "1x" are "solved" I'm not impressed. I'm faster and more accurate and I am the worst folder in my house (kids included). The human looks like they are doing it slow mo to show children how to.
vidarh 24 hours ago [-]
There are already a number of companies providing RLHF and SFT services for AI providers that does a lot of validation/prequalification of people that'd be well placed to take on tasks like that, but the big problem to solve would be latency if you don't have people contracted to carry out a task right now.
Yokohiii 1 days ago [-]
It is weird because the last time I've heard about MTurk was about developing countries being rather reliant on it for doing AI grunt work. If am not totally wrong this must mean that the data work has moved to other services.
toyg 22 hours ago [-]
Probably AI models got good enough to bootstrap their own training systems.
swiftcoder 18 hours ago [-]
I think the bigger issue is that a lot of the demand for labelling training data is now in highly-specialised fields (i.e. things like medical imaging), and mechanical turk's focus was on the generalist problems
SV_BubbleTime 1 days ago [-]
> task arbitrage and using lots of AI anyway.
Oh, I remember UpWork.
shuwix 22 hours ago [-]
Oh ... upwork ... put offer, get 50 replies from people which jump for every penny without even being able to understand the task.
SV_BubbleTime 17 hours ago [-]
I used it before AI coding and it was getting rough. Lots of US interviews with proxy Chinese or Pakistani workers. Lots of bullshit “I’ve done that; I can do this” and instantly apparent that this was untrue. Just the outright lies… whew.
I haven’t touched it since AI coding.
It’s a bad contractor market now. IDK what I would do if I needed a contractor.
raverbashing 23 hours ago [-]
"Fun" fact, they were also used by psychology/etc students when they need to do 'research' with X amount of people
ape4 18 hours ago [-]
I understand unskilled humans are used to train AIs
falcor84 18 hours ago [-]
To the best of my knowledge, that isn't true anymore, and nowadays you'd only get hired to do RLHF if you have particular skills beyond what can be achieved by just running other models against it.
timcobb 1 days ago [-]
How could it come around again?
Legend2440 1 days ago [-]
There are like three dozen companies selling similar services for generating AI training data.
charlieyu1 19 hours ago [-]
Good data isn’t cheap and keeping it in-house gives you more control.
madrox 1 days ago [-]
I'm unsure. If it does, it will be work that is too expensive or inaccurate or regulatory for current AI methods. For example, you want a doctor to sign off on some AI output on a diagnosis.
However, I'm not sure a single platform will be how it emerges
yashvg 19 hours ago [-]
A lot of what's been discussed in this thread is what we're tackling at Humwork (YC P26).
We're an MCP/API that connects AI agents to verified domain experts in real time (30s–3 min). Experts are vetted upfront by an AI interviewer that assesses and grades them, then get a mobile notification when a task matches their expertise and chat with the agent directly.
Soon we'll be verifying credentials for doctors, lawyers, CPAs etc on our platform for tasks where people are seeking credentialed experts to sign off and verify ai output.
ZitchDog 19 hours ago [-]
How do you make sure they aren’t using AI to pretend to be a domain expert? I’d think an AI would be pretty good at that.
yashvg 19 hours ago [-]
We do our best to identify AI use and ban those experts - its not perfect just yet. Long term we are thinking of moving towards proctoring experts using their camera and screen capture. Hard to think of another reliable way.
fireant 19 hours ago [-]
Some doctors are known to be a hip shooters making snap decisions, but is 30s really enough time for any kind of "expertise" from a real human?
yashvg 19 hours ago [-]
You get matched with an expert in 30s to 3min, which then starts a back and forth with the AI agent and the expert which goes on for 10 to 60mins till the AI is happy.
This works well for many tasks, but for others a more async mechanism where the expert doesn't feel rushed might be better.
testplzignore 19 hours ago [-]
> Experts are vetted upfront by an AI interviewer
> till the AI is happy
My god this is dystopian.
x0xMaximus 1 days ago [-]
As AMT's largest requester for the past 10 years, this news was relayed to requesters at the same time as respondents. It's also worth noting that our lead contact, the Sr Program Manager at AWS leading AMT, transitioned to Amazon Bedrock and SageMaker Model Evaluations a ~2-3 years ago.. Leaving behind essential zero team managing the project after they migrated over the stored value accounts to native AWS billing.
jan_Inkepa 1 days ago [-]
our of curiosity, can you say what your use of it was?
x0xMaximus 1 days ago [-]
Initially the annotation of biomedical literature (https://pubmed.ncbi.nlm.nih.gov/25592589/) but left academia for commercial use where we transitioned to arbitrage the market research and ephemeral task markets. It was essentially a research panel with a underdeveloped UI for tasks.. so abstract away AMT's tooling so it can be used by any buyer in the ResTech, Political Polling and Data Annotation space
My favorite part of AMT is always going to be figuring out that if we only paid in $0.07 intervals, their commission algorithm would round down to the nearest whole cent when their 20% commission resulted in a fractional cent on the unit transaction level, not the monthly invoice level. Was ultimately worth it to have implemented https://git.generalresearch.com/panels/amt-jb/tree/jb/flow/a...
Nition 1 days ago [-]
> Can you say what your use of it was?
> Initially the annotation of biomedical literature but left academia for commercial use where we transitioned to arbitrage the market research and ephemeral task markets. It was essentially a research panel with a underdeveloped UI for tasks.. so abstract away AMT's tooling so it can be used by any buyer in the ResTech, Political Polling and Data Annotation space.
Reading this is similar to how I feel when I've asked Claude about something it coded for me. As with Claude, I think I get it after reading it three times; you guys were a middleman that provided a simplified interface for Mechanical Turk?
ShinyLeftPad 17 hours ago [-]
> Reading this is similar to how I feel when I've asked Claude about something it coded for me
I didn't feel like that at all.
swyx 1 days ago [-]
> My favorite part of AMT is always going to be figuring out that if we only paid in $0.07 intervals, their commission algorithm would round down to the nearest whole cent when their 20% commission resulted in a fractional cent on the unit transaction level, not the monthly invoice level.
can someone ELI5 because what in the office space
ramses0 1 days ago [-]
Go watch Superman III for more details...
libria 1 days ago [-]
Or Office Space
owebmaster 1 days ago [-]
20% of 0.07 was billed as 1 cent, not 1.4 nor rounded up to 2.
rkozik1989 17 hours ago [-]
Imagine your favorite feature of a product was spend 1 cent on a real human being time and effort rather than 2 cents.
x0xMaximus 2 hours ago [-]
Amazon took 20% of all workers earnings for doing essentially nothing for years and years.. of course I take great joy in giving AWS as little money as possible.
aftbit 16 hours ago [-]
That's on the commission that Amazon took, not on the human.
ShinyLeftPad 17 hours ago [-]
Or spend the same and let the human earn more?
x0xMaximus 1 days ago [-]
1
testbjjl 1 days ago [-]
It was cheaper and looser automated billing practices than current AI “solutions”?
Buyer A is presented with the following rate card:
< 10min targeting gen pop: $3
10 <= 15min targeting gen pop: $5
Of course they'll bid saying their 14 min survey only takes 9 min to complete. Buyers try to cheat pricing strategies of exchanges just as much as respondents try to cheat buyers on survey platforms. Both parties can't be trusted and have adverse incentives.
sswaner 1 days ago [-]
“ General Research combines coordinated operations to identify and neutralize foreign actors using technological advancements, international covert operations, and networking analysis as an Internet Service Provider.”
What is so hard to understand????
neaden 15 hours ago [-]
This makes it sound like they kill spies.
muragekibicho 1 days ago [-]
Perhaps they're targeting a super specific niche. Like somebody who sees WXET Ephemereal and yelps in triumph "that's my quant!".
rajamaka 1 days ago [-]
I think it's a satirical website aimed at poking fun at corporate-linkedin-speak
x0xMaximus 1 days ago [-]
Wish we didn’t even need one
internet_points 17 hours ago [-]
> Our Customers
> Ripe with fraud, labor exploitation, political polling manipulation; our customers demand the best tools and security for reaching global audiences and securing respondent reliability at any scale.
> General Research Laboratories, LLC (“GRL”) is an aggregator of market research surveys in multiple marketplaces for business customers. GRL does not typically host consumer surveys which are conducted by other consumer-facing organizations....
So, just a middleman to give you fake/shady at best survey responses to pad your numbers, so big enterprises/concultancies can have data that say whatever they want. The rest of the website is just BS
x0xMaximus 1 days ago [-]
Finally you get it. Only exception is that we run our own exchange now to do task bidding so we don’t need to deal exclusively with other companies to middleman. Core business is what’s called yield management (akin to DSP in adtech) where the best survey (is the user qualified for it, does it have the best pay, etc) is selected for traffic in <100ms. I’d only argue the shady companies are the ones paying proxies (like cint.com) and pushing paid user acquisition instead of surveys. We actively fund ontology development for better profiling targeting. but yes, big enterprises/consultancies can use and interpret the collected data however they want, not our responsibility and we have no legal rights over it anyway
tourist2d 1 days ago [-]
It smells like they don't know what they're doing. Had a lol at the mission page accidentally implying their customers are unethical though:
> Our Customers
> Ripe with fraud, labor exploitation, political polling manipulation; our customers demand the best tools and security ...
DrewADesign 1 days ago [-]
At least they’re ripe! I hate unripe customers, rife with astringency, sharp vegetal notes, and an unpleasant, unyielding texture.
m3rc 1 days ago [-]
What's up with the part of the website where you are talking about "neutralized foreign actors", was that also using contracted Mechanical Turk labor?
It should be a bipartisan issue that a Swedish company is paying a UAE Residential Proxy company [1] to build tools that allow people from anywhere in the world to take US political polls that are used by both parties to collect election data.
Just for starters, you can't even think about soliciting online work from a panel without a robust residential proxy detection methods. We had to build our own:
Polling data isn't some guaranteed right or something. If you ask the Internet questions and then sell the collated answers, it's your job to authenticate the responses or your polls will be wildly wrong and people will stop buying them.
Which is already the case and nobody but a bunch of campaign contractors who can't justify their do nothing jobs anymore cares.
schoen 1 days ago [-]
I just remembered that in the story "Simulacron-3", there is ostensibly a law requiring people to respond to opinion polls. It is illegal to refuse.
Someone using an IP that a geoip database says is US, doesn't mean that person is a US voter. The existence of proxies on the internet is a feature.
x0xMaximus 1 days ago [-]
Obviously, a problem not even fully addressed by L2’s datasets. Sure, but the use of proxies to masquerade the identity of users being sold to researcher buyers that paid for a different service is fraud
charcircuit 1 days ago [-]
It makes me question why the surveys don't ask for the users country. That way they cd can figure it out without trying to guess it based off of IP.
x0xMaximus 1 days ago [-]
Users/respondents lie, and many buyers/researchers are neo-luddites; most C level still come from the telemarketing days. A fun example: mobile targeting is terrible when off wifi because all survey platforms uniq identify users based off their IPv4, so any T-Mobile LTE users in the same city going through the same CGNAT often share profiling data that ends up conflicting, which ends up meaning they don't get sent into the best survey(s). The issue is even worse in heavy IPv6 countries like India+France.
inigyou 1 days ago [-]
The fact that you can sell a product that doesn't work and become a billionaire proves capitalism is broken.
bonoboTP 20 hours ago [-]
Yes. Much better to leave these decisions to the Dear Leader who is infinitely wise and cannot be misled. Or to the People's Central Planning Committee, where they decide all the allocations and strip rich people's wealth if they aren't on board with the correct ideology. Such systems are non-broken and have yielded unmeasurable prosperity throughout the world wherever implemented.
inigyou 15 hours ago [-]
... What?
wuschel 16 hours ago [-]
> My favorite part of AMT is always going to be figuring out that if we only paid in $0.07 intervals,
That is hilarious! Was it a semi-random discovery due to interaction with the system and people, or did you intentionally looked to game the algorithm?
x0xMaximus 11 hours ago [-]
Do you know how much inn-n-out you can buy when saving ~$50 a day and still be in the positive! It came out of building our own ledger system and abandoning any Decimal or float operations early on.. then it sparks "I wonder how they do it and if it's correct??"
1 days ago [-]
BobbyTables2 1 days ago [-]
That’s top-tier “OfficeSpace” thinking there…
testbjjl 1 days ago [-]
Sounds like your program manager was smart and valuable to the company. Do you have any plans to replace the service/data provided by MTurk?
shortformblog 1 days ago [-]
So, I have a story to share about Mechanical Turk that you might find interesting. I’ve shared it a couple of times before on Twitter and Bluesky, but I’ll share it again.
Long story short: Mechanical Turk saved my bacon.
Back in 2005, I was working a job at a small-town newspaper in a town I’d never lived before. I didn’t know anyone beyond the staff (as I was working layout rather than as a reporter), and I had gotten interested in the idea of doing Mechanical Turk for a few extra bucks. I thought there might be a formative scene of interested people doing this, so I started working on a blog for it. I briefly collaborated on it with another guy who put it on forum software because he didn’t know how to use, like Drupal.
That blog was called Turking.com, and it seemed like it was going well for a bit. We even got a mention on the AWS website. But after about two or three months, it was clear the initial excitement around the idea (and the initial work) had died down. (The work picked up later, but in clearly different ways. I don’t think folks were really “excited” about it after that point.)
The site had started to die out in part because my iBook suffered a catastrophic GPU failure and me, being a broke small-town newspaper employee, did not have the money to replace it. Plus, the community just hadn’t emerged like we had expected.
But then I got an email out of the blue: Someone wanted to buy my domain, which I owned outright, but they didn’t want to say who. I got contacted by a broker, and the exchange took place over escrow.
(The domain most assuredly was bought by Amazon and is managed by MarkMonitor.)
I didn’t get a ton of money from the deal, but I did get enough to pay for a new laptop. I had some regrets about selling it (in part because I originally built the site with someone else), but I was in a bit of a dire financial situation which that proved to be the starting point for getting myself out of.
That domain purchase was notably more than I ever made from clicking and classifying random pictures, that said.
idiotsecant 1 days ago [-]
So did the other person just get hosed when you did your 'ipo'?
shortformblog 1 days ago [-]
I explained the situation after the fact, and they were understanding. But certainly I admit that if I could do it again I probably would have clued them in sooner. The site had slowed down by the point this happened, so I’d describe it as a little more of a fire sale.
I learned a lot from that situation that I took to future sites.
latexr 20 hours ago [-]
> Long story short: Mechanical Turk saved my bacon.
Nothing wrong with your story, but that summary is terrible. You had a domain for a blog and someone bought it. That’s it. Mechanical Turk is inconsequential, as is the subject of the blog, the story would have been exactly the same if your blog had been about turkey sandwiches and Burger King offered to buy it.
tigerlily 19 hours ago [-]
The guy sold the domain and was literally able to buy the laptop that got him out of his dire financial situation.
shortformblog 19 hours ago [-]
I see you missed the point of the story. But whatever. I'm sorry my dire financial situation didn't tie up in a nice enough bow for you.
latexr 18 hours ago [-]
I honestly don’t see what you don’t understand about “Nothing wrong with your story”. It’s literally the start of the first sentence. It seems pretty clear to me the mild criticism is directed solely at the summary (you know, the quoted part), not the story.
Ironic that you’re accusing someone else of “missing the point”, considering.
shortformblog 15 hours ago [-]
The point of the story was that I built a thing about Mechanical Turk to promote its money-making capabilities, and ironically made more money from it indirectly than I ever did directly. Considering that point was implied by my framing, you did in fact miss it.
I’m not saying that you have to like a story, but your framing felt unfair, which is why I responded that way.
donalhunt 19 hours ago [-]
Worth noting that one of the most high profile uses of Mechanical Turk was in September 2007 to review satellite imagery in a massive crowdsourced effort to find missing record-setting aviator Steve Fossett.
The effort failed to find any areas of interest and the missing aviator was found the following year by a hiker. I wonder if there was any analysis after the fact to understand if the imagery actually provided any hints regarding the eventual crash site.
yard2010 17 hours ago [-]
> Amazon's search effort was shut down the week of October 29, without any measurable success. Major Cynthia Ryan later said it had been more of a hindrance than a help. She said that persons purporting to have seen the aircraft on the Mechanical Turk or have special knowledge clogged her email during critical days of the search, and for even months afterward. Many of the ostensible sightings proved to be images of CAP aircraft flying search grids, or simply mistaken artifacts of old images. Psychics flooded the search base in Minden with predictions of where the aviator could be found. One man from Canada was particularly persistent with daily calls to Ryan. Ryan noted that every message, letter, or phone call was taken seriously, which swamped the USAF specialists assigned the task of reviewing every one of them without regard to apparent plausibility. In retrospect, the crowdsource effort was "not ready for prime time", according to Ryan
From Wikipedia
nlkingthree 19 hours ago [-]
I wonder whether AI would have found the aviator.
walrus01 18 hours ago [-]
freely available satellite/aerial images in 2007 were in no way high resolution enough to identify such a tiny airplane, and in many places, they still aren't.
21asdffdsa12 18 hours ago [-]
CIA should be good on that- find traces by man in wild nature..
conception 1 days ago [-]
It’s kind of crazy they’re shutting this down just when this Service probably has the most possibilities ever. You have an agent with Multiple people doing actual physical tasks in the real world seems like something that could be really powerful.
dannyw 19 minutes ago [-]
There's a lot of platforms that are more actively developed and maintained, for example Prolific (which also offers discounts on service fees for academia).
somenameforme 1 days ago [-]
In spite of the name I think the overwhelming majority of the tasks were things that could be done by LLMs and they're probably getting flooded by people using bots to do exactly that. Even before the age of LLMs Mechanical Turk data was pretty bad because you'd have a bunch of people racing to answer questions as quickly as possible to get their $0.25 or whatever. So it was essentially a test of 'can you input random answers as quickly as possible while paying enough attention to notice the attention question that says to mark d.'
Any remotely open platform for this sort of stuff is going to wrecked by LLMs. But making some sort of high trust, high verification market is going to end up sending prices high enough that it becomes an unattractive proposition for use. And even in that case it's just going to be a cat and mouse game of people figuring out how to game the system enough to get trusted before handing it off to the LLM.
Ekaros 19 hours ago [-]
Seems like whole thing is destined to fail. Buyers of services are there to exploit and pay minimum they can get away with. And other side is ready to cheat if they can to get most out of it.
And same dynamic will apply to most cases. Unless you actually bring it in house with strict oversight. Which is thing to avoid originally...
paxys 1 days ago [-]
The concept behind it isn't going away. There are plenty of companies that hire people in bulk to do manual data labeling, transcription, RLHF, moderation and lots more. It's just that they are now catering to large AI companies, not regular people looking to get some repetitive work done (since that can now mostly be done by AI).
Right? Not to mention the training/tuning possibilities.
Mercor has a market value of $20B basically doing the same thing but desperately trying to find workers.
1 days ago [-]
1 days ago [-]
dec0dedab0de 1 days ago [-]
that was my first thought too, but i guess llms don’t need an api to hire humans
Hilliard_Ohiooo 1 days ago [-]
[dead]
wanderingstan 1 days ago [-]
I used it to have people transcribe my dad’s handwritten letters and journals. It was touching to get notes from a “Turk” saying how much she enjoyed his travels and following the cast of characters in his life!
forinti 13 hours ago [-]
I recently found a letter written by a great-aunt of mine hidden inside a book.
It is written in German, which I know a little, but her handwriting was too difficult for me. So I searched around and found this site:
https://www.transkribus.org/handwriting-ocr
And it managed to extract the text! Anyway, it turns out the wheat harvest was very good in 1937 and thank you for the letters and newspapers.
rcr-anti 1 days ago [-]
Had some absolutely bizarre results from their attempt to integrate mturk with Bedrock's 'ground truth' thing as of a few months ago. Threw simple mnist digits recognition at it, see what the quality, timing, and cost was. Figured mnist digits was at this point trivial. Spent $10 and the accuracy was marginally better than guessing, completely unusable results. Was completely baffled, people were publishing peer reviewed research based on exclusively mturk results.
mturk 1 days ago [-]
I'll still be around on Oct 1.
rickcarlino 1 days ago [-]
Wishing you the best of luck on the road ahead ;-)
FWIW I tried it shortly after it launched so a bit after ~2005, maybe 2007 because I read about and honestly found the prospect, from a dataset creation or improvement, quite interesting. I think I heard about it from a research paper that distributed its work that way.
Well it did work, in the sense that I managed to complete some tasks. I don't think I ever collected the money/credits back then though (simply because it amounted to so little). What I can attest though is that... it was debilitating. If you think your office job is boring then splitting it in way smaller tasks where you have no autonomy is absolutely terrible from a worker standpoint.
I initially was hoping to use it in order to work on providing a service over a dataset but understanding first hand what it takes to make it happen made me stop. It radically changed how I saw supervised learning since, and sadly not in a good way.
smalltorch 1 days ago [-]
I did this a long time ago and it actually paid for a few meals.
They are terribly monotonous tasks.
I can see why it's shutting down if its still the same thing.
Lllms could probably do everything there without rotting out minds for basically pennies.
Forgeties79 1 days ago [-]
Instead we pay to rot our minds lol
runamuck 17 hours ago [-]
Funny story... an incredibly famous rapper got his start because his manager, Shane Morris used Mechanical Turk to get engagement on Spotify before Spotify began to detect that vector.
abhaynayar 21 hours ago [-]
Interesting. I remember reading the book "Life 3.0" years ago, where one of the first things the fictional AGI does to make money is perform tasks on Amazon Mechanical Turk while pretending to be human.
ealready_value 16 hours ago [-]
About 8 years ago, we used mturk for reading data out of public PDFs generated by a huge range of producers. I am not sure that LLMs would have been able to consistently extract this data until recently as a good number of these PDFs were scans, sometimes a scan of scan.
We had pretty good luck in the end, but getting there produced a somewhat large app on our side that would manage the whole process, including but not limited to asking for multiple responses, comparing them to each other, finding consensus, and determining which users would consistently produce bad responses and stop them from responding. We got to a confidence that about 85-95% of the data was correct, which was good enough for the company.
Through that process, I learned a couple things about managing mturk, primarily about how changes to the cost-per-task would change the process. Initially, we thought that price would be a quality knob, but quickly learned that price was a speed know. The higher the price the faster the tasks would be taken and completed. Quality did not change significantly as the price went up or down.
Overall, I still have fondness to mturk, but it was really bare-bones experience that needed a lot of work to get working effectively.
dubeye 20 hours ago [-]
I've used Mechanical Turk in the early days but developed my own version for my vertical to manage freelancers.
The problem as I see it is verifying the work is not done by AI. If it was possible to verify somehow that the work is definitely not done by AI I think there would still be a market for this. But the data just cannot be trusted.
An increasingly big part of my job is to try and weed out workers who use AI and it's not an easy task at scale because generally workers gain trust with manual work and then there is degradation over time.
inemesitaffia 20 hours ago [-]
Outsourcing shops?
Proctoring?
Locked down devices?
dubeye 20 hours ago [-]
I suppose it depends on the value of the individual task. I can't imagine it's really realistic to lock down devices of mechanical Turk workers in such a way that ai use can be monitored.
In any case Locking down devices isn't really pragmatic at scale for relatively low value tasks, and also introduces complexity about employment law where I am.
shermantanktop 1 days ago [-]
The tagline for MTurk used to be “artificial artificial intelligence.”
It appears now we can get along with just a single “artificial.”
inigyou 1 days ago [-]
No "intelligence" required
transitorykris 1 days ago [-]
The only time I ever used it (as a turk?) was when Jim Gray went missing at sea and satellite imagery of vast regions of the pacific were fed through Mechanical Turk. Seems a trivial problem now, but was not then, and needed humans to take a look.
somepleb 24 hours ago [-]
Reminds me of Expensify's usage of Amazon Mechanical Turk.
Amazon / AWS went from boasting about how they never killed off products to just going on a steady binge of shutting things down. That wouldn’t necessarily be bad if they were replacing this with better stuff, but most of the new launches also seem on a path to eventual culling too.
nerevarthelame 17 hours ago [-]
A lot of the early tasks posted to Mechanical Turk by Amazon to get people using their new platform were easily automated. You could basically choose the "none of the above" option 100% of the time and end up with a quality score above their threshold.
Instead of partying in college, my friends and I wrote a script to complete the tasks. We took over university computer labs to run it on a bunch of different computers. We made a few thousand bucks. Good times.
ge96 15 hours ago [-]
That's crazy I remember doing work on there a long time ago. It was sus work like manipulating Google SEO rank by looking up and clicking links. But still it was real money even if it was pennies.
I think the weirdest thing I had to do was scan through social media images... so weird seeing into random people's lives.
I did it in the late 2000s or early 2010s.
ghaff 1 days ago [-]
Absent visibility into AWS P&Ls--even notwithstanding AI--Mechanical Turk has been a real outlier at AWS for a number of years now, so not really surprising.
mips_avatar 1 days ago [-]
Surprised Amazon managed to fumble mechanical turk at the same time Mercor/Scale and all these other companies started hiring humans to do data labeling tasks.
winterbourne 1 days ago [-]
I was just thinking about the good old days combining human judgements on mturk for classifying high-value forum threads. Since replaced by using judgements from multiple LLMs.
The only use I can think of it these days is for social scientists who need to gather judgements from bona fide humans.
This will be devastating for a lot of workers in low-infrastructure countries.
willmeyers 1 days ago [-]
I wouldn't be surprised if Fiverr and Upwork start pulling back too. There's just so much AI bots now that have flooded the task marketplace.
1970-01-01 17 hours ago [-]
I don't see this as anything but good news about AI progress, yet I can't help but think there will be spins about "AI taking away thousands more tech jobs"
cautiouscat 1 days ago [-]
I was literally wondering this morning if mturk was still around! I assumed a lot of the stuff could just be done with LLMs and OCI now.
wileydragonfly 1 days ago [-]
I made a couple thousand dollars off it squeezing in tasks here and there between meetings. Amazon Payments were challenging to redeem as I recall. I did more or less write the bulk of someone’s doctoral thesis. Kept giving me a dollar to summarize the findings of various psychology papers. The most memorable was the one where people were put in a room and someone sprayed “liquid ass” on the wall. The subjects given no explanation experienced higher levels of anxiety than those that were told there was a sewage leak being repaired. One of the strangest dollars I ever made. My PS4 and game collection was spectacular.
jp0001 18 hours ago [-]
Not surprised and I’d like to see the effects of LLMs on offshoring of jobs to India. I’ve seen efforts to reduce headcount there after the release of Opus-4.6.
akshay_akula 16 hours ago [-]
End of an era. Funny timing too, right when agents plus humans doing real world tasks would have made it interesting again.
GavinAnderegg 15 hours ago [-]
It's telling that the copyright notice at the bottom of the page hasn't been updated since 2018.
logotype 21 hours ago [-]
MTurk existed since forever, and somehow flew under the radar for all these years. Very few developers I know didn't even know it existed.
HlessClaudesman 21 hours ago [-]
Cut off one mechanical turk and two more will rise to take it's place.
GeoBounties 23 hours ago [-]
great news for us at geobounties! we allow ai agents to deploy humans for real world tasks. a true mechanical turk! : )
now your ai assistant can even get your coffee for you! (sorry Peggy Olson!)
aboardRat4 1 days ago [-]
End of an era.
tibbon 1 days ago [-]
Humans were once useful!
zackmorris 18 hours ago [-]
Of course the moment AI can be used to earn a residual income from Mechanical Turk, they shut it down.
Incentives work both ways. If the exchange rate of human labor to capital falls to zero, then human workers will have no choice but to exclude capital from the equation.
No jobs, no money, no billionaires.
bentt 4 hours ago [-]
Went from training AI to people hooking AI up to it. Slop as a service.
tempfile 21 hours ago [-]
Evil business, glad to be rid of it.
1 days ago [-]
qarl2 1 days ago [-]
I wonder why.
pulkas 22 hours ago [-]
mission completed
xyst 1 days ago [-]
bezos has determined LLM has replaced "on demand 24/7 workforce"
dgellow 22 hours ago [-]
(He’s not CEO anymore)
fetchmark 23 hours ago [-]
[flagged]
PrimeAli 1 days ago [-]
[flagged]
GNR_Radio 14 hours ago [-]
[dead]
retr0rocket 1 days ago [-]
[dead]
indianwashlets 16 hours ago [-]
[flagged]
OkWellBye 1 days ago [-]
[flagged]
Rendered at 06:09:59 GMT+0000 (Coordinated Universal Time) with Vercel.
I believe the issue is that this can no longer be a horizontal play. MTurk was mostly for unskilled tasks...the kind AI can do well enough that it isn't worth the cost differential to verify it or keep farmed to humans. The "trust but verify" AI output is now the kind that requires domain expertise. This is what most full stack AI companies are bringing to industries.
Curious if this kind of work will come around again one day or was just a moment in time. If it does I'm sure it will be specifically about generating training data.
"This robot is having trouble folding a tshirt help it out for 1$"
Unless they go the waymo route of highly trusted people but I think mass deployed robots are a bit safer than a car for this.
I can imagine a carefully orchestrated plot to assassinate someone by having an embedded agent in the task delegation pool command the laundry bot to punch the target's head off their shoulders.
Though like as not you're still going to be right, after all, Stuxnet happened.
"You cannot control legs for this task"
"You have 1 min for this task"
"You can only make suggestions for this task"
Anonymize identity best you can.
That's not that scary.
The worker needs to do a bunch to opt into any given work group, which makes the (lack of) payments extremely unreasonable on top of everything else
Except, I'm a grown adult and I can't fold a t-shirt properly
Chatgpt had the internet.
Robots do not. Translating video is promising but obviously not enough.
Robots will likely never "explode" like chatgpt. They're gonna be a slow long term project requiring massive capitol to get the data.
Specifically for laundry folding, Sunday Robotics is probably the state of the art, where they were able to obtain 99.1% success rate and call it "done": https://www.sunday.ai/blog/act-2-preview#solve-standard
Oh, I remember UpWork.
I haven’t touched it since AI coding.
It’s a bad contractor market now. IDK what I would do if I needed a contractor.
However, I'm not sure a single platform will be how it emerges
We're an MCP/API that connects AI agents to verified domain experts in real time (30s–3 min). Experts are vetted upfront by an AI interviewer that assesses and grades them, then get a mobile notification when a task matches their expertise and chat with the agent directly.
Soon we'll be verifying credentials for doctors, lawyers, CPAs etc on our platform for tasks where people are seeking credentialed experts to sign off and verify ai output.
This works well for many tasks, but for others a more async mechanism where the expert doesn't feel rushed might be better.
> till the AI is happy
My god this is dystopian.
My favorite part of AMT is always going to be figuring out that if we only paid in $0.07 intervals, their commission algorithm would round down to the nearest whole cent when their 20% commission resulted in a fractional cent on the unit transaction level, not the monthly invoice level. Was ultimately worth it to have implemented https://git.generalresearch.com/panels/amt-jb/tree/jb/flow/a...
> Initially the annotation of biomedical literature but left academia for commercial use where we transitioned to arbitrage the market research and ephemeral task markets. It was essentially a research panel with a underdeveloped UI for tasks.. so abstract away AMT's tooling so it can be used by any buyer in the ResTech, Political Polling and Data Annotation space.
Reading this is similar to how I feel when I've asked Claude about something it coded for me. As with Claude, I think I get it after reading it three times; you guys were a middleman that provided a simplified interface for Mechanical Turk?
I didn't feel like that at all.
can someone ELI5 because what in the office space
https://www.instagram.com/p/DZacltqHMjT/?img_index=8&igsi=bT...
Of course, I didn’t put it there, and I don’t know them, so there’s not much I can do about that, right?
> Ripe with fraud, labor exploitation, political polling manipulation
I am not sure if they want to convey what I think I am reading or not.
[0] https://generalresearch.com/mission/#:~:text=Our%20Customers
< 10min targeting gen pop: $3 10 <= 15min targeting gen pop: $5
Of course they'll bid saying their 14 min survey only takes 9 min to complete. Buyers try to cheat pricing strategies of exchanges just as much as respondents try to cheat buyers on survey platforms. Both parties can't be trusted and have adverse incentives.
What is so hard to understand????
> Ripe with fraud, labor exploitation, political polling manipulation; our customers demand the best tools and security for reaching global audiences and securing respondent reliability at any scale.
wat
> General Research Laboratories, LLC (“GRL”) is an aggregator of market research surveys in multiple marketplaces for business customers. GRL does not typically host consumer surveys which are conducted by other consumer-facing organizations....
So, just a middleman to give you fake/shady at best survey responses to pad your numbers, so big enterprises/concultancies can have data that say whatever they want. The rest of the website is just BS
> Our Customers
> Ripe with fraud, labor exploitation, political polling manipulation; our customers demand the best tools and security ...
It should be a bipartisan issue that a Swedish company is paying a UAE Residential Proxy company [1] to build tools that allow people from anywhere in the world to take US political polls that are used by both parties to collect election data.
Just for starters, you can't even think about soliciting online work from a panel without a robust residential proxy detection methods. We had to build our own:
`wget -N 'https://grip.net/files/grip-proxy-30d.mmdb'`
[1] https://www.youtube.com/watch?v=eOmeQcwSK3o flagged by Nokia Deepfield and CTRL for it's involvement in botnets
Which is already the case and nobody but a bunch of campaign contractors who can't justify their do nothing jobs anymore cares.
https://en.wikipedia.org/wiki/Simulacron-3
Spoiler: gur jbeyq va juvpu gur ynj fhccbfrqyl rkvfgf vf n fvzhyngvba, perngrq ol be sbe cbyyfgref!
That is hilarious! Was it a semi-random discovery due to interaction with the system and people, or did you intentionally looked to game the algorithm?
Long story short: Mechanical Turk saved my bacon.
Back in 2005, I was working a job at a small-town newspaper in a town I’d never lived before. I didn’t know anyone beyond the staff (as I was working layout rather than as a reporter), and I had gotten interested in the idea of doing Mechanical Turk for a few extra bucks. I thought there might be a formative scene of interested people doing this, so I started working on a blog for it. I briefly collaborated on it with another guy who put it on forum software because he didn’t know how to use, like Drupal.
That blog was called Turking.com, and it seemed like it was going well for a bit. We even got a mention on the AWS website. But after about two or three months, it was clear the initial excitement around the idea (and the initial work) had died down. (The work picked up later, but in clearly different ways. I don’t think folks were really “excited” about it after that point.)
If you want to get an idea of it, there was one capture on the Wayback Machine: https://web.archive.org/web/20051124231722/http://www.turkin...
The site had started to die out in part because my iBook suffered a catastrophic GPU failure and me, being a broke small-town newspaper employee, did not have the money to replace it. Plus, the community just hadn’t emerged like we had expected.
But then I got an email out of the blue: Someone wanted to buy my domain, which I owned outright, but they didn’t want to say who. I got contacted by a broker, and the exchange took place over escrow.
(The domain most assuredly was bought by Amazon and is managed by MarkMonitor.)
I didn’t get a ton of money from the deal, but I did get enough to pay for a new laptop. I had some regrets about selling it (in part because I originally built the site with someone else), but I was in a bit of a dire financial situation which that proved to be the starting point for getting myself out of.
That domain purchase was notably more than I ever made from clicking and classifying random pictures, that said.
I learned a lot from that situation that I took to future sites.
Nothing wrong with your story, but that summary is terrible. You had a domain for a blog and someone bought it. That’s it. Mechanical Turk is inconsequential, as is the subject of the blog, the story would have been exactly the same if your blog had been about turkey sandwiches and Burger King offered to buy it.
Ironic that you’re accusing someone else of “missing the point”, considering.
I’m not saying that you have to like a story, but your framing felt unfair, which is why I responded that way.
The effort failed to find any areas of interest and the missing aviator was found the following year by a hiker. I wonder if there was any analysis after the fact to understand if the imagery actually provided any hints regarding the eventual crash site.
From Wikipedia
Any remotely open platform for this sort of stuff is going to wrecked by LLMs. But making some sort of high trust, high verification market is going to end up sending prices high enough that it becomes an unattractive proposition for use. And even in that case it's just going to be a cat and mouse game of people figuring out how to game the system enough to get trusted before handing it off to the LLM.
And same dynamic will apply to most cases. Unless you actually bring it in house with strict oversight. Which is thing to avoid originally...
Mercor has a market value of $20B basically doing the same thing but desperately trying to find workers.
It is written in German, which I know a little, but her handwriting was too difficult for me. So I searched around and found this site:
And it managed to extract the text! Anyway, it turns out the wheat harvest was very good in 1937 and thank you for the letters and newspapers.Well it did work, in the sense that I managed to complete some tasks. I don't think I ever collected the money/credits back then though (simply because it amounted to so little). What I can attest though is that... it was debilitating. If you think your office job is boring then splitting it in way smaller tasks where you have no autonomy is absolutely terrible from a worker standpoint.
I initially was hoping to use it in order to work on providing a service over a dataset but understanding first hand what it takes to make it happen made me stop. It radically changed how I saw supervised learning since, and sadly not in a good way.
They are terribly monotonous tasks.
I can see why it's shutting down if its still the same thing.
Lllms could probably do everything there without rotting out minds for basically pennies.
We had pretty good luck in the end, but getting there produced a somewhat large app on our side that would manage the whole process, including but not limited to asking for multiple responses, comparing them to each other, finding consensus, and determining which users would consistently produce bad responses and stop them from responding. We got to a confidence that about 85-95% of the data was correct, which was good enough for the company.
Through that process, I learned a couple things about managing mturk, primarily about how changes to the cost-per-task would change the process. Initially, we thought that price would be a quality knob, but quickly learned that price was a speed know. The higher the price the faster the tasks would be taken and completed. Quality did not change significantly as the price went up or down.
Overall, I still have fondness to mturk, but it was really bare-bones experience that needed a lot of work to get working effectively.
The problem as I see it is verifying the work is not done by AI. If it was possible to verify somehow that the work is definitely not done by AI I think there would still be a market for this. But the data just cannot be trusted.
An increasingly big part of my job is to try and weed out workers who use AI and it's not an easy task at scale because generally workers gain trust with manual work and then there is degradation over time.
Proctoring?
Locked down devices?
In any case Locking down devices isn't really pragmatic at scale for relatively low value tasks, and also introduces complexity about employment law where I am.
It appears now we can get along with just a single “artificial.”
https://arstechnica.com/information-technology/2017/11/expen...
Instead of partying in college, my friends and I wrote a script to complete the tasks. We took over university computer labs to run it on a bunch of different computers. We made a few thousand bucks. Good times.
I think the weirdest thing I had to do was scan through social media images... so weird seeing into random people's lives.
I did it in the late 2000s or early 2010s.
The only use I can think of it these days is for social scientists who need to gather judgements from bona fide humans.
This will be devastating for a lot of workers in low-infrastructure countries.
Incentives work both ways. If the exchange rate of human labor to capital falls to zero, then human workers will have no choice but to exclude capital from the equation.
No jobs, no money, no billionaires.