Documents: OpenAI is asking contractors to upload their work from current or previous jobs to evaluate its models, leaving it to them to scrub confidential info
To prepare AI agents for office work, the company is asking contractors to upload projects from past jobs …
Context & Ripple Effects
OpenAI’s contractor request extends its effort to evaluate models against real office work, but assigns the first pass at confidentiality to contributors. That approach sits uneasily beside a prior report that confidential AI-training materials were broadly accessible through shared document links.
The company has also reportedly tightened internal protections for sensitive model information through security changes including isolated sensitive data. The contrast makes the handling and provenance of externally supplied work central to confidence in workplace-agent evaluation.
First-order effects
- Contractors must decide what confidential material to remove before submitting current or former workplace projects, creating a new screening burden and potential exposure point for them and their employers.
- OpenAI gains access to more realistic work artifacts for evaluating office-oriented agents, while relying on contributor-led redaction to limit sensitive-data intake.
Second-order effects
- Organizations whose materials may be represented in submissions may tighten employee policies, contractual controls, and internal document access around external AI evaluation.
- Evaluation-data vendors and model developers face pressure to show stronger provenance, access, and redaction controls, particularly after reports of weak access controls around training documents.
Third-order effects
- If real-work evaluation becomes standard for workplace agents, competitive advantage will increasingly depend on governed access to representative workflows rather than only broad public-data benchmarks.
- The gap between strict protection of model IP and contractor-managed filtering of workplace data could draw sustained scrutiny over accountability for AI data handling; the earlier FTC records demand on model risks and a security incident shows that such practices can become a regulatory focus.
The trend: Workplace-agent development is shifting toward evaluations grounded in real operational artifacts, making data provenance and confidentiality controls a core product constraint.