You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Social Security Benefits are still adjudicated using occupational descriptions from the Dictionary of Occupational Titles (DOT), which was last updated in 1991. Because this is over 30 years old, it would be useful to measure how closely the DOT aligns with modern occupational data, like O*NET and the Occupational Requirements Survey (ORS).
Using these datasets, and the official DOT → O*NET crosswalk, I plan to measure which occupations have changed the most since DOT was updated, what kinds of changes have occurred, and if those changes can affect current allocation of Social Security Disability Insurance, and Supplemental Security Income (the column we care about here is “Blind or Disabled”).
This is important for a few reasons:
These programs distribute hundred of billions of dollars annually, so we obviously want eligibility decisions to be based on accurate information.
In 2024, Social Security stopped allowing 114 DOT occupations from use in disability decisions because the occupations barely existed. Many more occupations are likely outdated.
Technological change is accelerating, particularly with AI. If occupational information is decades out of date, rapid technological change could further widen that gap.
What are this project's goals? How will you achieve them?
I want to measure which occupations have changed the most since DOT, measure patterns in that change, and measure if that change can affect disability benefit allocation.
A rough pipeline would be
Crosswalk DOT occupations to modern occupations using official DOT/O*NET crosswalk to get the corresponding DOT and O*NET codes/titles. This is fully deterministic and needs no LLM judgement.
Use a three model LLM judge ensemble to compare DOT descriptions against O*NET tasks and activities. Three frontier models will receive the same records and score how much the content of the occupation has changed.
Measure changes in job requirements with ORS. ORS gives information that O*NET does not provide cleanliness, like the distribution of lifting requirements, time spent standing, etc. This can be performed mechanically.
Run extensive quality and sensitivity checks. Test questionable mappings, model agreement, prompt sensitivity, etc.
Aggregate the results, estimate how many occupations are basically unchanged, materially changed, transformed, or no longer have a modern counterpart, and identify what kinds of occupations changed most.
Map the results to Social Security disability adjudication, and test whether using modern occupational information changes which jobs claimants may be considered capable of performing.
A potential extension is combining the measured historical changes with an AI occupational exposure score to identify which occupations may change the most/fastest in the future.
The funding will pay for the LLM judging in Step 2, including multiple independent frontier model judges, disagreement adjudication, reruns, robustness checks, and sensitivity analyses. The exact models will depend on funding.
Solo authored. A bit about me:
I’m an Emergent Ventures Fellow, my website yourjobrisk.com was on Marginal Revolution (Tyler Cowen’s blog).
My prior project on Manifund has been submitted to a journal, and will be available as a SSRN preprint soon.
I coauthored an occupational exposure measure, which is still in progress
The most likely cause of this project failing is that DOT is close to modern measures (based on my brief preliminary I think this is not the case). If so, that is an interesting finding in itself
The second risk is the reliability of LLMs as judges. I have used similar methods successfully before, but it still needs to be handled carefully. I’ll mitigate risk by using independent models, running a pilot before the full experiment, testing agreement and prompt sensitivity, and retaining all outputs.
There may also be mapping problems - a lot of DOT occupations map to broader occupational categories. I will measure and test sensitivity to this.
$20,000 from Emergent Ventures, and $2,750 from my previous Manifund project.
There are no bids on this project.