Can LLMs help build a psychological scale?
A Saudi study used several AI tools to help draft and refine a 17-item emotion-regulation measure. Results in 743 adults are encouraging but preliminary, local and partly post-hoc—not clinical validation.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1Researchers used several LLMs and AI tools to assist item generation, refinement, bias screening, translation and interpretation, reducing an initial 25-item pool to a 17-item Emotion Regulation Scale under human oversight.
- 2A convenience sample of 743 adults was divided into non-overlapping exploratory and confirmatory subsamples; a four-factor solution explained 50.85% of variance and the post-hoc refined model showed acceptable initial fit.
- 3The cross-sectional local sample, post-hoc refinement and absence of independent replication, test-retest evidence or predictive validation mean the scale is not ready for diagnosis or high-stakes individual decisions.
Research topic
Whether multiple LLMs can assist item generation and refinement within a conventional human-led psychological scale-development process
The answer: LLMs can assist the drafting process, but they did not validate the instrument
A peer-reviewed study from King Abdulaziz University offers an unusually concrete example of using multiple large language models inside a conventional psychometric workflow. AI tools helped generate, refine, screen and translate candidate items for a new Emotion Regulation Scale, while the researchers retained human oversight. In an initial convenience sample of 743 adults, the resulting 17-item measure produced a four-factor structure and several expected correlations with established mental-health and wellbeing measures.
That is encouraging early development evidence, not permission to deploy the scale for diagnosis, treatment selection, hiring or other high-stakes individual decisions. The sample was cross-sectional and locally recruited, model refinement was partly post-hoc, and the paper reports no independent replication, test-retest stability or evidence that scores predict future outcomes. It also does not isolate whether using LLMs produced a better scale than expert-only development would have done. The contribution is a documented assisted workflow and a candidate instrument that now needs tougher human validation.[1]
The Impact Brief · Free
Follow the evidence in society & media.
Get a five-minute weekday briefing on what changed, why it matters and where the evidence comes from. Choose the topics you care about.
How AI was used—and where human responsibility remained
The researchers grounded the measure in Gross's Process Model of emotion regulation rather than asking a chatbot to define a construct from scratch. Multiple LLMs and other AI-based tools assisted with item generation, wording refinement, bias screening, translation and interpretation after exploratory factor analysis. The starting pool contained 25 items; the final scale contained 17. The paper describes human oversight throughout, which matters because item meaning, cultural appropriateness and the decision to retain or remove an item are scientific judgements rather than language-generation tasks.
The workflow nevertheless creates a new reproducibility problem. A future team needs to know which model versions, prompts, language settings and decision rules shaped the candidate pool, because commercial systems can change silently. Agreement among several models is not independent expert validation; systems may share training material and linguistic conventions. Bias screening by an LLM can flag awkward wording, but it cannot prove an item functions equivalently across sex, age, language, disability or cultural groups. Those claims require participant data and formal measurement-invariance tests.[1]
What was tested in 743 adults
The cross-sectional study recruited 743 adults through convenience sampling. Participants were randomly divided into two non-overlapping subsamples: one for exploratory factor analysis and one for confirmatory factor analysis. That separation is stronger than exploring and confirming a structure in exactly the same observations, because the confirmatory group has not directly determined the initial factor extraction. The report does not turn the convenience sample into a population-representative one, however; who volunteers for an online psychological survey can differ from people reached through probability sampling or clinical services.
The exploratory analysis supported four preliminary factors and explained 50.85% of observed variance. The confirmatory analysis supported a post-hoc refined 17-item model with chi-square divided by degrees of freedom of 1.95, comparative fit index of 0.934, Tucker-Lewis index of 0.918, incremental fit index of 0.935 and root mean square error of approximation of 0.050. These are useful fit indicators, but post-hoc refinement can adapt a model to this particular dataset. A locked specification in a new sample is the more demanding next test.[1]
Reliability and correlations are adequate for a first study, not definitive
Internal-consistency estimates ranged from 0.68 to 0.74 for Cronbach's alpha and from 0.68 to 0.75 for omega. The authors describe those values as acceptable for early-stage development, an appropriately cautious framing. They show that items within a factor were related, but internal consistency does not prove that a scale is one-dimensional, stable over time or sensitive to meaningful change. A very high coefficient can also result from repetitive items, so reliability needs to be interpreted with content coverage rather than used as a stand-alone quality stamp.
The proposed adaptive dimensions correlated positively with cognitive reappraisal, with coefficients from 0.522 to 0.608; the maladaptive-cognition dimension correlated 0.400 with suppression. Reported concurrent associations also ran in expected directions with depression, anxiety, wellbeing and quality of life. For example, maladaptive cognition correlated 0.432 with depression and 0.462 with anxiety, while adaptive strategies showed smaller negative correlations with both. These cross-sectional relationships support plausibility, but cannot show direction, mechanism or clinical usefulness.[1]
What psychologists and service designers should do with this finding
For researchers, the study suggests that language models can expand and critique an item pool quickly, especially when bilingual phrasing or several conceptual domains must be considered. The safer pattern is the one the paper attempts: begin with a defined theory, document AI assistance, keep accountable human reviewers and subject every retained item to participant-based psychometric testing. LLM output should be treated like a set of suggestions from an uncredentialled assistant, not like respondent evidence or a substitute for people with lived experience.
For clinicians and public-service professionals, the current scale should remain a research instrument. A total or subscale score must not be converted into an automated judgement about a person's diagnosis, risk, eligibility or treatment. Before any service use, developers should establish norms, clinically meaningful thresholds, test-retest reliability, sensitivity to change, accessibility and performance across demographic and language groups. People completing a scale also need to know what the result means, who sees it, how long it is kept and whether a human can correct an inappropriate interpretation.[1]
Limits, independence and the evidence needed next
The authors reported no specific funding and declared no competing interests. Ethics approval came from the university's Faculty of Arts and Humanities, and participants gave electronic informed consent. The article is an accepted, peer-reviewed version released early and may still receive copy-editing before the final version of record. The study's main limits arise from design: one convenience sample, one cultural setting, self-report outcomes, cross-sectional correlations and a post-hoc refined confirmatory model. No comparison arm tested whether AI assistance improved development quality, time or bias relative to established expert practice.
The assessment would change with preregistered replication of the locked 17-item structure in independently recruited Saudi and international samples. Researchers should report test-retest stability, measurement invariance, responsiveness, missing-data behaviour and comparison with clinician or behavioural outcomes where appropriate. A transparent audit of prompts, model versions and human changes would make the AI contribution reproducible. Direct comparison with an expert-only workflow could test whether LLM assistance adds coverage or efficiency without importing stereotypes. Until then, the paper shows feasibility and preliminary validity—not that AI can manufacture trustworthy psychological measurement on demand.[1]
What this means for people
- Psychologists may gain a faster way to expand and review candidate items, but remain responsible for theory, wording and validation.
- People completing mental-health measures need human explanation and safeguards against scores becoming automated labels.
- Services should not use this preliminary scale for diagnosis, eligibility or treatment decisions without independent validation.
Global context
The work comes from Saudi Arabia and broadens a literature often dominated by North American and European samples. That regional contribution is valuable, but it does not by itself establish cross-cultural equivalence. Emotion-regulation language, help-seeking and response styles vary, so translation and LLM fluency cannot replace local cognitive interviewing, representative sampling and measurement-invariance analysis in every intended population.
What the evidence does not yet show
- The 743 adults were convenience-sampled in one national setting, limiting population and cross-cultural generalisation.
- The study was cross-sectional, so correlations cannot establish direction, causation, stability or prediction.
- The confirmatory model was refined post hoc and has not been locked and replicated in an independent sample.
- No expert-only comparison isolated the effect of LLM assistance on item quality, bias, time or cost.
- Test-retest reliability, measurement invariance, predictive validity and clinical decision thresholds were not established.
What to watch next
- Independent preregistered replication of the locked four-factor, 17-item structure.
- Measurement invariance across languages, cultures, ages, sexes and relevant clinical groups.
- Test-retest reliability, sensitivity to change and links to external behavioural or clinical outcomes.
- Transparent model-and-prompt records and direct comparison with an expert-only scale-development process.
Living evidence record
Impact record IAI-1MRJX9N
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
11 October 2026
Source trail
1 direct source across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Single-source reporting disclosure
This record analyses one direct source. It can establish what BMC Psychology published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.
Evidence trail
Sources used for this report
Links checked 11 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Society & Media
When should newsrooms disclose AI use?
A peer-reviewed case study based on 13 interviews with 12 Financial Times managers and 28 internal documents finds that AI disclosure is treated as a spectrum shaped by oversight, risk and context. It describes one newsroom's practice and does not test whether labels improve audience trust.
7 min · 1 source
Society & Media
Can an inaccurate AI summary change what you remember seeing?
In a US online experiment with 328 analysed participants, accurate recall of a traffic sign fell from 83.6% after a consistent summary to 44.8% after a misleading one. The study isolates a classic misinformation effect; it does not measure real police reports or prove that ordinary model errors cause the same size of harm.
7 min · 2 sources
Society & Media
Does moderate AI involvement feel more creative?
Across two Chinese online experiments with 600 mostly student participants, a workflow preserving meaningful human choices produced higher self-rated creativity than low- or high-involvement alternatives. The tasks were short, the workflows bundled several features and the result is not evidence that the finished art was objectively better.
7 min · 1 source
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.