Normal view

There are new articles available, click to refresh the page.
Yesterday — 18 September 2026Main stream

Chrome Can Now Measure Exactly How Obnoxious Your Website’s Ads Are

17 September 2026 at 15:00
Chrome can now measure exactly how obnoxious a website’s ads are, from how much of your screen they swallow to how much data and CPU time they burn. Those ad-stuffed websites can finally be measured, and the results are public.
Before yesterdayMain stream

What I learned in my summer internship researching digital content accessibility

11 September 2026 at 14:30

In this post, User Experience team intern Hannah Watson shares her work over the summer researching digital content accessibility with the EdWeb 2 publishing community.

Introduction

As part of my internship with the User Experience Service, I have been investigating digital accessibility at the content level across University web pages, particularly on EdWeb 2 sites. Digital accessibility at the content level is about applying principles of accessibility to how web content is written and formatted. This includes, but is not limited to, content features such as heading levels, links, and alt text. It does not include anything that is not involved with content design or that is controlled at a higher level by the Content Management System (CMS), such as font, font size, or colour contrast.

Research aims and scope

The specific research questions for this project were:

  • How do web publishers learn about digital content accessibility?
  • What do web publishers know about digital content accessibility?
  • How do web publishers implement digital accessibility requirements and principles in their content?
  • What challenges do web publishers face in creating digitally accessible content?

More comprehensively, I was investigating the accessibility of:

  • Heading levels
    • Ensuring that heading levels are used correctly
    • Not using heading levels for emphasis
  • Links
    • Clear and concise link text which describes the linked destination
    • Link text which makes sense on its own
    • Avoiding URLs on web pages
  • Lists
  • Alt text
    • Clear link text that describes the image
  • Images
    • Ensuring that all images have appropriate alt text
    • Avoiding images of text
  • Videos
    • Including human-corrected captions with any videos uploaded to web pages
    • Making sure that transcripts are available for all videos uploaded to web pages
  • Italic, bold, and underlined text

For more detailed guidance on these topics, please refer to the University of Edinburgh editorial style guide.

Editorial style guide | Information Services

This research has been necessary for the User Experience team as it allows us to identify which areas of content accessibility are challenging for web publishers. From this, the team can adapt guidance and training to provide extra support on these more challenging areas where possible.

Research methods

Interviews with web publishers

I started my research by setting up short, informal interviews with six University web publishers. In these interviews, I asked the publishers about their experiences of creating digitally accessible content. The aim of these conversations was to understand how publishers learned about digital accessibility and what challenges they face when making content as accessible as it can be. During the interviews, we referred to web pages that these publishers work on to get concrete examples that illustrate the topics we discussed. Through these interviews, I was able to identify a number of trends, particularly in the challenges that the interviewees and their colleagues face.

Survey of EdWeb 2 publishers and the Web Accessibility Special Interest Group

Following these interviews, I created a survey to further investigate the findings and collate more supporting evidence for these findings. The survey consisted of 10 questions, three of which were demographic based as a filter, with a final question asking permission to follow up with those who responded.

The other six questions asked about:

  • which actions the participants took to make their content digitally accessible
  • what resources they used to do so
  • what challenges they face in doing so

The findings from the survey were effective in making the information gathered in interviews more robust and provided further evidence for some trends that were identified previously.

To publicise this survey, I sent a brief statement explaining my research into two Teams channels, the Web Accessibility Special Interest Group and the EdWeb 2 Community. Overall, the survey received five responses, and while this is a limited number, I found that it helped to support the findings from the interviews.

Analysis of Effective Digital Content workbooks

The final method of research that I used to learn about the digital accessibility of content on University web pages was by looking at pages that were submitted within Effective Digital Content workbooks as part of the course. The course requires learners to choose pages from a University website and assess the effectiveness of the content. By looking at the pages that learners selected, I was able to use active examples and make note of content accessibility issues that were present on live pages. This process also highlighted a number of trends.

This method of research was also useful in that it provided two separate sources of data. Firstly, the answers in the workbook helped me to gauge learners’ understanding of the principles covered in the course. Secondly, the web pages linked by those taking the course allowed me to have live examples of content to assess against content accessibility principles. While there was not necessarily overlap between the answers in the workbook and the pages submitted as part of the workbook (as some people may not have been fully or at all responsible for the content on those pages, or the content could have been updated since the workbook was submitted), it was useful to see these things separately.

Findings about accessibility

Combining findings from all areas of research for this project, I have identified a number of trends in the accessibility of digital content.

Headings, alt text, and links are the principles most often put into practice by participants

In the interviews and the survey, participants were asked what principles of accessible digital content they actively used when designing content. More than half of participants mentioned three principles in particular, with all participants mentioning at least two, which were:

  • using the correct heading levels
  • adding meaningful alt text to an image
  • writing clear and descriptive link text

Interestingly, while these principles were mentioned frequently by participants, and evidenced by the websites that we discussed in interviews, they are also principles that are often missed on University web pages. This is discussed in more detail in the corresponding sections further on in this post.

PDFs are hard to avoid

One standout finding from the interviews was that web publishers sometimes struggle to find an effective alternative to PDFs. While PDFs are not necessarily accessible, they do have benefits which make them useful for publishers. For example, they are downloadable, searchable, and cannot be easily edited without permission from the owner. There are ways to make PDFs more accessible, such as avoiding decorative images, adhering to accessible content design principles within the PDF, and checking colour contrast. However, the interviewees were more in favour of finding a way to turn their content into a webpage as this is more likely to result in accessible content.

Out of the survey responses, three also mentioned that they had recently chosen to publish content as a web page rather than as a PDF, highlighting their knowledge that a web page is preferable in terms of accessibility. However, their personal preferences or how difficult they found this is unknown as I was unable to follow up with these participants.

Heading levels are often skipped and headings vague

Across web pages that I assessed for correct heading levels, there were several which skipped heading levels throughout the page content, such as going straight to a heading 3 without that heading being nested within a heading 2 section. Additionally, headings on the University web pages that I investigated were often generic, instead of being specific about what a page or section will contain, which is the recommended approach.

The fact that the web publishers who were interviewed and surveyed were aware of the importance of correct headings levels and specific headings, and that the majority of these publishers have attended a staff training or used the editorial style guide, suggests that the guidance provided is accurate and useful for publishers.

To increase the use of correct heading levels, as well as clear and descriptive headings on web pages, participant responses suggest that increasing the reach and engagement of existing training and resources involving headings would be effective. This includes training provided by the User Experience Service, such as Effective Digital Content or Content Improvement Club, and the University of Edinburgh editorial style guide.

Alt text is well written but sometimes missed

In three of the interviews, and in two survey responses, participants mentioned that they struggled to find the time to add alt text to images on their web pages. However, participants in the interviews also stated that they understood the importance of alt text, and what writing meaningful and clear alt text involves. This is reflected in the answers in Effective Digital Content workbooks. The course contains a question asking learners to write meaningful alt text for two images, and this question is often answered well. This suggests that web publishers understand how to write alt text, and that the obstacle in doing so is more likely to be related to time and resource.

One way that time constraints on adding alt text could be improved is to reduce the number of images on a web page, which will not only make this task more manageable but also is also more sustainable.

Explaining acronyms and abbreviations is common

All five responses to the survey stated that they had explained an acronym or abbreviation recently. Although this did not come up in the interviews often – only once – the survey responses suggest that this is common practice for web publishers who are familiar with digital accessibility principles.

Link text is frequently inline or not descriptive

Participants stated that they understand the importance of clear link text as a principle and make effort to implement this into their content. However, similar to headings, this sentiment is not reflected across numerous University web pages. In some of the pages submitted as part of the Effective Digital Content workbooks that I investigated, there were frequent occurrences of inline link text, or link text that does not clearly describe the linked destination.

To increase the writing of link text on a separate line, as well as clear and descriptive link text, participant responses suggest that increasing the reach and engagement of existing training and resources involving links would be effective. This includes training provided by the User Experience Service, such as Effective Digital Content or Content Improvement Club, and the University of Edinburgh editorial style guide.

Findings about staff experience and engagement

A trend from both the interviews and survey responses is that for many of the people involved with this research, their knowledge of digital accessibility started with a personal interest. Specific examples of this that participants mentioned include learning about accessibility as a student (and then going on to be an accessibility advocate for a student society) and working with disabled students and learning through experience.

During the interviews, multiple staff members mentioned that they believe, through their experiences, that a large part of issues with creating accessible digital content at the University surrounds communication. A combination of factors was discussed that had communication at the centre, including:

  • the importance of digital accessibility not being widespread enough.
  • consistency about expectations between schools or areas of the University, such as one page being edited by multiple schools or areas and having different standards.

Job-based constraints were also mentioned frequently, such as limited time to add alt text to all images on a web page and working on a page where the lead publisher takes a design forward approach which can sometimes clash with accessible content principles. These answers were also reflected in the survey responses, with all of the responses mentioning either one or both of these problems.

What resources do web publishers use for learning about and developing their digital accessibility skills?

The primary resource that participants mentioned using to learn about digital accessibility and develop their skills was University-provided staff training, with 10 of 11 participants mentioning this. Effective Digital Content was specifically mentioned twice, and Content Improvement Club three times. In the interviews, two participants mentioned more general staff training, with one survey response saying the same. In the survey responses, three participants also said that they used training provided by the Disability Information Team.

The Web Content Accessibility Guidelines 2.2 (WCAG) is another resource that was frequently mentioned by participants, with six in total saying that this is a resource that they use as guidance on accessibility.

The University of Edinburgh editorial style guide was mentioned by five participants as a resource that they used to provide guidance on digital accessibility, in which the guidance reflects what publishers learn in training courses, meaning that information they take from the style guide is in line with accessibility and content design training.

What I learned from researching content accessibility approaches in EdWeb

I learned a lot during my time researching how web publishers at the University of Edinburgh approach creating accessible digital content. I thoroughly enjoyed the opportunity to work with staff from a variety of different areas of the University. This helped me to learn how to identify commonalities in interview and survey responses despite the areas of work being distinct. I also enjoyed this because I was able to learn much more about the work that goes on across the University and what the work of other teams involved.

I also particularly enjoyed the format of my internship being a combination of individual work and working with other members of my team and the Disability Information team. This balance allowed me to set my own goals while still getting the opportunity to work and learn collaboratively as part of a team, prioritising my tasks between my own and those that I was working on with others.

A limitation of my work was that it was much easier to contact and work with staff who have a genuine interest in the subject area of accessibility, which does not lead to research that is representative of how University staff approach digital accessibility as a whole. Having done the research that I have so far, a continuation of the research would be most beneficial if it focused on the experiences and approaches of staff who are less familiar with digital accessibility requirements and principles, as this would create a more well-rounded understanding of web publisher accessibility approaches at the content level.

Furthermore, another limitation of the work was the time constraints which have restricted my ability to develop solutions to issues identified during the research period. This was expected to a degree, as the solutions depended on the outcomes of the research, which has taken the majority of the 12-week period. The positive outcome of this is that the research and findings will provide the User Experience service with information that allows for both further research and for adaptations to guidance and training if it is deemed necessary.

Conclusion

This research has helped the User Experience Service to assess their training and the resources that are available for publishers to learn more about digital content accessibility. Keeping in touch with the publishing community through projects like this helps the team to direct their effort to real challenges that publishers face on a day to day basis.

The research completed throughout the duration of this project will be supported and advanced by further investigation, particularly by communicating with web publishers who are less involved with digital accessibility as a whole. This will likely help in providing more concrete solutions to digital content accessibility issues across EdWeb 2 pages.

Why Tech Companies Hid the “Log Out” Button (And How to Win It Back)

1 September 2026 at 12:33
Ever noticed that signing in takes one tap, but logging feels like solving a puzzle?That's not bad design. It's a business strategy. The modern internet has quietly turned "leave" into the hardest button to find—and most people never realized it was happening.

The Infinite Scroll Was a Design Crime. Now Meta is Paying for It.

25 August 2026 at 11:30
The infinite scroll was never a UX feature—it was a slot machine disguised as code. Now, as Meta faces a historic trial over algorithmic addiction, web designers must finally confront the ugly truth: we sacrificed basic usability and human willpower just to keep users trapped in a bottomless pit.

Hearing about approaches to student voice at the Learning and Teaching Conference 2026

In June, I headed over to the Nucleus Building to attend the University’s Learning and Teaching Conference. The theme of the conference was student voice, and I was there to present about how the Careers Service and the UX Service had involved students in the design process for a recent website development project.

Emily McIntosh’s keynote “Nothing about us without us”

The day kicked off with a keynote from Emily McIntosh, Chief Student Officer at Harper Adams University. Emily set up the theme of student voice and covered some important distinctions for us to keep in mind throughout the day. Emily compared two ways that students are involved in decision-making processes: representation and partnership.

Representation is where staff within a university create something – a new policy, for example – then show student reps and ask for their opinion. While this is better than not involving students at all, it can lead to student reps feeling that they are responding to decisions that have already been made. The input from student reps sometimes comes at late stage, and can lead to tweaks but not necessarily substantive change.

Emily suggested that “you said, we did” posters can be indicative of this sort of approach. She contrasted this with examples of universities working in partnership with students from the beginning. In this way, students are actively involved in helping to shape services and policy. Emily described this as a relational, dialogic culture of ongoing collaboration, where students and staff work on problems together. This approach is sometimes described as co-design.

For me, there were some clear analogies in what Emily was describing to processes we advocate for in user-centred design. For example, bringing user input into the early stages of a design process is generally more effective than getting the same input later in the day. In addition to this, user-centred design often emphasises the benefit of ongoing collaboration and continuous iterative improvement, keeping users involved before and after the launch of a product.

Working in partnership like this is not necessarily a simple thing to achieve. But I appreciated the message: incorporating students into how we decide to do things as a university brings some significant advantages. Whatever the practical challenges of implementing this approach, moving from designing-for to designing-with offers us opportunities to better meet service users’ needs.

Highlights from the rest of the day

The rest of the day made me think about issues around student voice.

Neil Speirs (Widening Participation Manager) raised the issue of working class students often being absent from student voice projects. Neil’s talk made me think about how good quality user research often depends on recruiting a diverse group of participants. Recruiting the right research participants can be difficult. We see that all the time in our work. But without a diverse range of voices, user research can miss important perspectives and draw an incomplete picture of what matters to users of a product or service.

Neil Speirs speaking in a lecture theatre in front of a slide titled A neutral proposition

Neil Speirs presenting at the Learning and Teaching Conference. Slide reads: “A neutral proposition? We must be aware that student voice does not stand as a neutral proposition. Without democratic and participatory structures and cultures, student voice projects in learning and teaching will not provide democratic and participatory outcomes. We might reflect on the class based inequalities that are ‘baked into’ (Major et al, 2023) our structures, cultures, policy, curriculum and pedagogy. Which students are readily able to participate and benefit from student voice projects? Which students are invited to take part? Do we even notice? Is it the usual suspects? Do all students know the rules of the game and therefore why and how it might be advantageous to partake in student voice projects? Is the working class student voice heard in these spaces?”

Clare McKay (Student Recruitment and Admissions) and Christine Emmerson (Communications and Marketing) talked about how student recruitment markets can help us to anticipate student needs. One important point they made was that marketing is not just about promoting courses and programmes. It’s also about understanding our audience and whether there is a need. This resonated with aspects of UX work, where an understanding of user needs can tell us not just how to communicate information, but also whether that information is needed in the first place.

Jenna Fyfe and Carine Abraham (Edinburgh Medical School) talked about research they had done with alumni networks. Their project focused on alumni from three online postgraduate programmes at the Medical School. Alumni described how they create networks with each other following graduation, and how they often retain an interest in maintaining contact with programme teams from the University. This contrasts with a picture where alumni wish to maintain contact with university-wide structures. I found this interesting, because we sometimes think of people outside the University seeing us as a single entity unified by an overarching brand. In contrast, the alumni in this research weren’t thinking about a monolithic institution. Instead, they retained an interest in personal connections with a small, specific part of the University.

My talk was about our work with the Careers Service

In my talk, I gave an overview of how the Careers Service developed their website in 2025. In particular, I highlighted how the User Experience Service assisted with the design of collaborative activities that brought student voice into the work. I wanted to show how bringing student input into this sort of design activity makes the work better, and I talked about the benefits that the Careers Service achieved by approaching the project in this way.

You can read more about our collaboration with the Careers Service elsewhere on this blog:

Blog posts about work by the UX Service and the Careers Service

Reflections on the day

I thought there was a lot of value in holding a conference like this, with most of the speakers being staff and students from the University of Edinburgh. Presenters were speaking from a broadly similar context, which meant the talks had a high level of relevance to our work. There were a lot of familiar faces and it was good to catch up in the breaks with colleagues I hadn’t seen for a while.

My main takeaway was that student voice projects often have a lot of crossover with user-centred design approaches. In both cases, we’re interested in making things better by involving end users. I felt it was a great opportunity to hear about how student voice projects have been informing colleagues’ work at the University.

About the conference

University of Edinburgh Learning & Teaching Conference

Can AI do a content audit of my website? A review of different ELM models and AI products

Content audits are a long-time staple of the user-centred website toolkit, but they take time and effort to complete, which many website owners struggle to find. In my ongoing AI experimentation, I tried using AI tools to help me with auditing web content.

Every day brings a raft of new AI developments. When choosing how to use AI, I’m less sold on getting it to do creative tasks for me because I like doing those myself. Instead, in pursuit of freeing up my time, I prefer to use AI to help me with the tedious tasks I never get around to, the ones I know will take me longer than I think, the ones I put off again and again. Here’s looking at you, content auditing.

Websites are easy to grow but difficult to keep in check

When people come to the UX Service for help with their websites, they tend to use several phrases to describe their site. These have included:  ‘It’s a mess!’ ‘It’s a bit out of control’. ‘It’s grown arms and legs’. None of this is unusual when you consider the lifecycle of websites. Over time, websites change hands. Pages once carefully crafted are easily forgotten as new editors take over. New content is created without knowing what’s already there. Very often, content needs to be published to a deadline and it’s quicker to publish a new page than seek out and amend an existing one. Seldomly do publishers want to get rid content that’s taken time and effort to produce.

Read some thoughts on content housekeeping by UX Service team members (current and past):

Digital housekeeping – applying content management practices to improve digital sustainability by me

Be a gardener by Ari Cass-Maran

Content Audit Findings and the 100k Challenge by Milo McLaughlin

How to get a grip of your website (and then keep hold) by Neil Allison

A content audit is the best place to start improving a website (not the homepage)

When people seek to bring order to disarray on a website it’s difficult know where to start, so typically, they start on the homepage with ideas to change images or add new content. While this gives sites a superficial makeover, it doesn’t really help site audiences in search of content. Fewer and fewer website audiences are starting out on homepages, instead they’re parachuting directly into site content from Google search and increasingly, from AI summaries. What does this mean? It’s more important than ever to make sure your web content is up-to-date, accurate, useful and relevant for your website users. Here’s looking at you, content auditing.

Read more about getting your content ready for AI from a recent Content Improvement Club session:

Auditing a website needs a methodical approach (which is where AI can help)

If audits are so important, why are they so easy to put off? Short answer, they can take prolonged amounts of time and concentrated effort. The larger and more complex your site is, it’s likely that there will be more content to review, and the more difficult it may be to maintain a consistent auditing approach. Going through multiple, long pages of web content, it can be easy to run out of the time you’ve allocated for your audit. Breaking the audit into chunks may lead to different results from different sessions as biases creep in. Bringing colleagues in to help can share the load, but can also introduce new perspectives and contradictory auditing decisions. There’s a need to maintain ruthless objectivity, to counter human subjectivity and avoid errors infiltrating. Here’s looking at you, AI.

If you can plan and describe a website content audit, you can ask AI to help you do it

Website content audits can be done for many purposes, but commonly, doing a content audit involves making decisions about what content to keep, what needs to go and what needs to be changed – all in pursuit of reaching a future state of your website. A good way to decide what content to keep is to consider who it’s for and the purposes it serves. If a piece of content doesn’t meet audience needs and doesn’t align with the site purpose, then it’s probably a candidate for deletion. A good content audit of a website therefore typically requires several things:

  1. An inventory of every piece of content in the site. Typically arranged in a spreadsheet, with one row per webpage. Can be organised into sections or content types, or grouped by root URL
  2. A note of the site audiences, ideally in order of priority (answering the question ‘Who is this site primarily for?)
  3. A list of the main reasons those audiences visit the site (answering the question ‘What tasks do people complete on this site?’) together with a list of the main goals of your site (answering the question ‘What do we want people to do on this site?’)

Considering each of these things as different pieces of data, where item 1 is data to be audited, items 2 and 3 are factors or auditing criteria used to decide what happens to those data (typically one of three outcomes: ‘Keep’, ‘Delete’ or ‘Modify’), I reasoned that I could give AI these data, tell it what a content audit was and and ask it to complete this process for a given website.

I experimented with ELM to help me do a content audit of a website

The University’s AI platform, ELM was the natural place to start experimenting with applying AI to aid content auditing. The ELM interface allows you to input a prompt, as well as add additional files and data sources. It also allows you to select different models to handle queries. I was curious to see the differences between the different models – both non-reasoning ones and reasoning ones so I needed to design a prompt that would work for both.

Read more about ELM and its models

ELM website

ELM new model guides (University log in required)

Being mindful of the environmental cost, I wanted to start with a small AI model (Llama 3.3)

AI comes with a significant environmental cost, which is highly dependent on the LLMs you use. Powerful models with more reasoning power are especially resource-intensive to run, so to avoid unnecessary wastage of tokens and other resources I’ve found it’s sometimes best to start with a smaller model and size up accordingly, as the need arises. The trade-off is that smaller models have limited reasoning power, however, so the responses they provide and the tasks they are able to complete can be limited.

Based on my previous experience of using AI and thinking about its potential application to a typical content auditing process, I had a hunch that providing a small model with something like a site inventory (such as a XML sitemap) and detailed data about website audiences and purposes could overload its context window and produce inferior results.

I therefore decided to begin with a lightweight prompt ‘Can you help me do a content audit of this website? (with a link to the website)’ to give a small model something manageable and learn how it approached handling the query. I chose a website that the UX team had recently worked on, so I was familiar with its content, and I entered the prompt, starting with the Llama 3.3 model – an open weights model (costed on computing power rather than per-token usage) which is hosted in University data centres (instead of being cloud-based). I toggled on the option to include web search so ELM was able to access the website.

The same prompt to different ELM models produced varied results

Having starting with Llama 3.3, I moved on to other models: Open AI’s GPT 5 and then Anthropic’s Sonnet 4.6. To assess and critique responses from each, I adopted the mindset of a person reluctant to begin a content audit and looking for AI help to kick-start action on this, and used this as a roleplay to review and compare the responses.

The response from Llama 3.3. was too generic to convince me to start auditing

Despite a glitch in formatting the response, Llama 3.3. provided some general information about content auditing. Its response included an initial assessment of the site, which made objective judgements of the homepage (assessing it as ‘well-structured, with a clear introduction’), the navigation menu (labelling it ‘easy to use with links to key sections’), content quality (noting it as ‘high-quality, informative and engaging’) and accessibility (pulling out that the site had an accessibility statement).

The response then acknowledged that in order to do a thorough content audit on the site in question there was a need to provide guidance on the aspects to focus on, and rounding off, the response noted some of the aspects a typical content audit would examine: content quantity, content quality, accessibility and consistency across different parts of a site.

Reviewing this response, I felt that if I was a site owner looking to improve my site, this response could very easily persuade me that I didn’t need to bother auditing it at all – and I wasn’t left any wiser about how to meaningfully provide guidance on the aspects to focus on if I did decide to proceed with an audit.

Likelihood to convince me to audit: 3/10

Open AI’s GPT 5’s response was detailed and thorough but a bit overwhelming

As would be expected from a reasoning model, the response when I used GPT 5 was more detailed than Llama 3.3’s. Arranged in two sections, the first part included preliminary observations using the site URL provided – which were split into the following categories:

  • Information architecture and navigation – commending the site’s positioning on user intents, yet suggesting improvements with more explicit audience pathways and call-to-action signposting
  • Content coverage and freshness – using the site’s range of content types to assess content structure and assessing freshness based on recency of news items
  • Call to actions – acknowledging that the site had were multiple calls-to-action which could be combined into primary CTAs for each audience
  • Accessibility and usability – calling out the need for more meaningful and less duplicated alt text on the site images
  • Trust and governance signals – recognising the role of the site’s strong branding to demonstrate authenticity
  • SEO – suggesting the site’s page titles and meta descriptions were reviewed for relevance

The second part of the response outlined a plan for a full audit. The plan began by posing three questions to be answered:

  1. What are the primary goals? e.g. increase visitors, boost traffic, showcase outputs
  2. Who are the priority audiences? (with suggestions of the different groups)
  3. What’s in scope? (suggesting either the entire site or specific sections)

The response then provided a suggested series of 7 steps to follow to complete the audit:

  1. Inventory and crawl (with detail of the inventory to start with and the tools to complete the crawl)
  2. Editorial quality review (with detail of criteria to score each page – such as inclusivity, accuracy, audience-fit etc)
  3. UX and IA review (with suggestions to evaluate navigation labels and connections between content and pathways to achieve top tasks)
  4. Accessibility (with suggestions of accessibility checking tools to use as well as manual checks to complete)
  5. SEO and technical (with suggestions to check data points like metatags, internal links, robots files and performance)
  6. Analytics (with suggestions to use analytic data such as bounce rates, page views, site searches etc)
  7. Recommendations and roadmap (detailing prioritisation measures to make the identified site changes)

It rounded off suggesting a template structure for the content inventory (and offering to create this template as a Google Sheet), an idea for a scoring rubric, and reiterating different tools to use (naming products like Screaming Frog, Google Analytics and SEMrush). To conclude, it offered to do the audit, requesting answers to the three initial questions and asking for access to a sitemap, external tools and governance documents (such as style guide and brand voice guidelines).

Reviewing the response, which was presented as text content, I found it a bit too much to scan and digest, and putting myself in the place of someone facing content auditing with some reluctance, it felt that this large amount of detail could be off-putting to even start the process. I could appreciate the relevance of suggesting of external tools, but knowing that each of those would require setting up accounts and logins, as well as going through a process to integrate with Open AI, it had the effect of making me mentally push content auditing to the bottom of a to-do list.

Likelihood to convince me to audit: 6/10

Anthropic’s Sonnet 4.6’s response gave structured guidance which was easier to follow

The output from the Anthropic model stood out compared to the others from ELM since it was not just plain text, instead, it had been structured with headings, tables and icons. This made scanning and digesting this response much easier than the others.

The response was structured into 10 sections, as follows:

  1. Site overview – table summarising the site name, owner, primary purposes and date of confirmed content
  2. Information architecture – list of the top-level navigation elements with an assessment of what worked well and issues identified in a bulleted list
  3. Content inventory – table of pages, with noted audiences, details of last update (if known) and inferred status (on a red, green and amber scale)
  4. Calls to action – list of CTAs on the homepage, with assessments and recommendations
  5. Content freshness – table containing list of most recent items in each section and related assessment
  6. Accessibility – list of accessible features that were present, table of items that needed checking and recommended accessibility tools to make the assessment
  7. SEO and technical – table containing an SEO assessment, considering elements like meta descriptions, structured data, mobile responsiveness – each with an allocated status, associated notes and combined recommendations
  8. Content quality and tone – list of observations and recommendations
  9. Priority recommendations summary – table of priority actions with an effort/impact score
  10. To complete the full audit I need – list of requirements to complete the full audit (including sitemap, CMS page list, Google Analytics data, etc.)

Content within each of these sections was comparable to that included in the GPT 5 response, but in the Sonnet 4.6 response, these sections referred directly to information from the site in question – for example, instead of making general assessments of site aspects like navigation and content coverage, it contained direct detail about the site itself with mention of specific pieces of content, noting what was already known (for example, referring to issues noted in the accessibility statement). Arranging this in tables with use of icons and red, green and amber priority indicators made it easy for me to pinpoint what to address and with what urgency, and having this consolidated in the final table was easier to parse than reading it in text format (as in the GPT 5’s response).

Reviewing this response from Sonnet 4.6, it was the best of the bunch from ELM, it contained enough detail to help me get started but not so much that it was overwhelming. Putting myself in the position of a time-poor site owner unsure where to start, I had clear pointers of the priority areas and a mix of quick wins and bigger tasks to choose from.

Likelihood to convince me to audit: 7/10

Three sets of two side-by-side screenshots, showing the output from Open AI GPT 5.5 (on the left) compared to the output from Anthropic's Sonnet 4.6 (on the right) showing Anthropic's more structured and visually appealing approach to presenting the responses

Three sets of two side-by-side screenshots, showing the output from Open AI GPT 5.5 (on the left) compared to the output from Anthropic’s Sonnet 4.6 (on the right) showing Anthropic’s more structured and visually appealing approach to presenting the responses

As well as ELM, I experimented with Chat GPT and Claude to provide content audit help

Models within ELM can also be accessed through interfaces provided by AI products. Two of the best-known AI products are Chat GPT (by Open AI) and Claude (by Anthropic). I was curious to see the responses each of these products would provide in response to my query asking for content auditing help so I prompted each of them in the same way. In both cases I used the free versions.

Chat GPT gave a suggested approach and produced initial findings

The response from Chat GPT was comparable to the output from ELM choosing an Open AI model, but the response was structured and annotated with tables and icons, in a style similar to the Anthropic response. It began with a suggested audit approach which set out different areas to consider and related questions. In the next section (headed ‘Initial observations’) it contained an analysis of the site itself, which included site strengths and highlighted the following six areas for improvement on the site in question:

  1. Homepage is quite academic – pulling out some audience questions that would be relevant to be answered on the homepage instead of the content that was there
  2. Navigation could be more task-focused – noting that the current navigation was driven by organisational structure rather than tasks users would want to complete
  3. News archive – picking out the need for recent articles
  4. Calls to action – acknowledging that many of these are worded to provide information not prompt an active response
  5. Content consistency – advising a uniform page structure for easier reading
  6. Audience segmentation – recommending clearer delineation between groups of site users

The response concluded with an example audit table and some content recommendations, organised under ‘Keep’, ‘Improve’ and ‘Review’.

Reviewing this response as a whole, I liked the up-front suggestion of an auditing approach and the example audit table which was then supplemented with detail about the site itself and recommendations relevant to different aspects of the site in question. This helped me visualise what the audit outputs could look like, and I could see ways to get there.

Likelihood to convince me to audit: 7/10

A screen showing the output from Open AI after being prompted for help content auditing. The suggested audit approach laid out in a table with columns for the different areas of auditing and associated questions to ask in the right-hand column

A screen showing the output from Open AI after being prompted for help content auditing. The suggested audit approach laid out in a table with columns for the different areas of auditing and associated questions to ask in the right-hand column

Another screen showing the output from Open AI after being prompted for help content auditing. An example audit table is laid out with rows for each URL, and columns containing associated purpose, audience, quality and recommendations.

Another screen showing the output from Open AI after being prompted for help content auditing. An example audit table is laid out with rows for each URL, and columns containing associated purpose, audience, quality and recommendations.

Another screen showing the output from Open AI after being prompted for help content auditing. Content recommendations are laid out with colour coding for 'Keep', 'Improve' and 'Review'.

Another screen showing the output from Open AI after being prompted for help content auditing. Content recommendations are laid out with colour coding for ‘Keep’, ‘Improve’ and ‘Review’.

Claude presented interactives to actively take me through an audit process

Anthropic’s Claude interface took a different approach to the other AI products, immediately asking about the main goal of the audit with options to pick. I picked ‘full inventory’. It then asked the desired breadth of the audit with another set of options, from which I picked ‘full site crawl’. The next screen asked me about the output format, and I picked ‘spreadsheet’.

From top to bottom, the interactive screens produced by Claude asking about the main goal of the audit, the breadth of the audit and the format for the audit output. Each has multiple choice answers to select

From top to bottom, the interactive screens produced by Claude asking about the main goal of the audit, the breadth of the audit and the format for the audit output. Each has multiple choice answers to select

After approximately 2-3 minutes, with a running commentary of what was being done, Claude presented a list of ‘headline findings’ of bugs, orphaned content, out-of-date content and broken content (dead links, incomplete calls-to-actions etc). It also produced a downloadable Google Sheet with two tabs – the first containing a summary of the issues to be fixed, their locations and a reason why they needed to be fixed, the second with a full inventory, with one row for every URL in the site and columns logging the page title, content type, date, status (with colour-coded options: ‘Keep’ (green), ‘Update’ (amber), ‘Rewrite’ (dark amber), ‘Review’ (blue) and ‘Critical’ (red)) and recommended actions for each page.

Screenshot of the audit output spreadsheet produced by Claude, showing one row per page URL and different columns logging content types, date, status and recommended actions.

Screenshot of the audit output spreadsheet produced by Claude, showing one row per page URL and different columns logging content types, date, status and recommended actions.

 

Reviewing this response, it was by far the most effective at making me feel I was starting the content auditing process and supporting me through it. Being presented with the opportunity to answer questions upfront meant I could control how the audit went, and having a spreadsheet created for me to review saved a lot of effort making this from scratch by scraping the site for a sitemap and page metadata.

Likelihood to convince me to audit: 9/10

Conclusion: If you’re putting off auditing your website, try getting AI to help

If you’re looking to improve your website, doing a content audit is one of the best steps you can take – working out what you have, what’s working and what’s not will put you in a strong place to move forward and make positive changes. Depending on how well you know your site, and how much time and resource you have to spend on it, you may appreciate some help to get started auditing its content.

From my quick experimentation with the different AI tools available I learned that, with the exception of the small non-reasoning model, each AI product offered something to help initiate the content auditing process of a website. At a basic level, asking AI to look at your website will pull out something you may not find yourself such as an out-of-date page, an irrelevant CTA or an accessibility glitch. AI can also take a zoomed-out look at your site that you may be too close to adopt yourself – such as reviewing your navigation pathways, appraising your audience segmentation or analysing the consistency of your content voice and tone.

If you’ve never audited a website before, AI can offer a lot of guidance to get you started, and whether you choose to do it yourself taking recommendations from the AI on tools and/or processes or whether you prefer to hand the heavy-lifting to AI to get your audit spreadsheet started – leaving you to deal with the more nuanced decisions that require specialist knowledge and experience, these tools can save you time and effort and may just be the nudge you need to do this all-important website maintenance task.

 

Using scenario-based research to understand how colleagues find inclusive language guidance

Earlier this year, the User Experience (UX) Service researched how colleagues find and use the University’s Inclusive Language Guide. This post explores why we used scenario-based research, what it revealed about colleagues’ behaviour and how the findings are improving the guide. 

The Inclusive Language Guide is only useful if people can find it 

The Inclusive Language Guide sits within the University’s Editorial Style Guide and provides practical guidance when writing about disability, race and ethnicity, and sex, sexuality and gender. The guide was developed in 2022 through a co-design process involving colleagues from across the University.

Inclusive Language Guide 

Having useful guidance is only part of the challenge. Colleagues also need to know it exists and be able to find it when they need it. 

More recently, we reviewed the Inclusive Language Guide to understand how discoverable it was for colleagues creating digital content. Emma Horrell has written about how this work prompted improvements to the guide’s visibility and signposting across the University. 

You don’t know what you don’t know: Improving the way we position inclusive language at the University

Why we used scenario-based research 

We wanted to understand more than whether colleagues were aware of the Inclusive Language Guide. Awareness alone does not tell us what people do when they need support, so we used realistic scenarios to observe how colleagues responded to situations where they might need guidance. 

We focused our research questions on four areas: 

  • How colleagues approached decisions about inclusive language. 
  • Where colleagues looked for support and whether they could find relevant guidance. 
  • Colleagues’ awareness of the Inclusive Language Guide and what they expected it to help with. 
  • Any gaps, ambiguities or areas of confusion. 

These questions shaped the structure of the interview and activities we designed. 

Researching colleagues approach to inclusive design decisions 

We combined semi-structured interviews with scenario-based activities to understand both colleagues’ existing experiences and how they approached real content decisions. 

We spoke to nine colleagues who create digital content 

We carried out nine interviews with colleagues in different roles across the University. Participants had a range of experiences creating digital content, including two student interns who brought perspectives of those newer to the University’s digital landscape. 

We began by exploring participants’ existing experiences 

Before looking at the scenarios or the guide itself, we asked participants about the types of content they create and whether they had ever needed to check terminology or wording before publishing. For example, we asked whether they had been in a position where they needed to consider how to word something to make sure the language was not excluding anyone. 

This helped us understand the approaches colleagues already used and the sources they relied on when dealing with these topics. 

Realistic scenarios showed how colleagues look for guidance 

We introduced a series of realistic scenarios based on everyday content tasks. 

Rather than asking whether participants knew about the guide, the scenarios allowed us to observe how they approached these situations in practice.

Example scenarios included: 

  • writing promotional content about a Pride initiative and wanting to ensure the language used was fair and respectful 
  • writing about someone who had experienced trauma and wanting to avoid misrepresenting them 
  • checking the preferred terminology for the chair of a company board before publishing a news article 

In posing the scenarios, we gained insight into where participants would usually look for this type of guidance, how confident they felt making language decisions and what they understood by the term ‘inclusive language’.

We adapted our scenarios as we learned more 

One of the original scenarios asked participants how they would write about someone elected to lead a company board. Rather than focusing on terminology, several participants described how they would research the individual or organisation before publishing. 

While this reflected realistic editorial practice, it did not help us answer the research question we were exploring. 

For later sessions, we refined the scenario so that it focused specifically on checking the preferred terminology for the role before publication. This encouraged participants to think about language choices and produced richer insights into where they expected to find guidance. 

What we learned: Four themes around finding and using inclusive language guidance

Across interviews, there were four themes that emerged consistently. 

Colleagues valued inclusive language even when they had not heard of the guide 

Across the interviews, participants consistently demonstrated that they valued using respectful and inclusive language. They were already thinking carefully about inclusive language, but they did not always know where to access University specific guidance.  

Many described that they would check terminology with colleagues or seek advice from specialist teams when writing about unfamiliar topics. They wanted confidence that they were using appropriate language. 

Awareness of the importance of inclusive language was much higher than awareness of the Inclusive Language Guide itself. 

Colleagues expected to find guidance through different routes 

The scenario-based activities helped us understand where colleagues naturally looked for support. 

Participants frequently described turning to colleagues, managers, Equality, Diversity and Inclusion (EDI) resources, staff networks and external websites. A few participants did eventually land on the guide, but they had not recalled it or looked there first. 

Several participants also relied on previous examples when writing content. Others spoke about balancing official guidance with lived experience and perspectives from the communities being represented. 

This highlighted an important consideration. Colleagues expected inclusive language guidance to be connected to the places they already look for support, linked from elsewhere rather than existing solely within the Editorial Style Guide.  

Improving signposting therefore remains just as important as the guide itself. 

Colleagues understood the phrase ‘inclusive language’ in different ways 

Participants interpreted the phrase ‘inclusive language’ in different ways. 

Some described it as “using the correct terminology”, while others framed it more broadly as “not leaving anyone out”. Those with accessibility backgrounds often drew direct connections between inclusive language and accessibility. 

Despite these differences, the term ‘inclusive’ was recognisable, with several participants using it unprompted when working through the scenarios.

Colleagues looked for practical guidance first 

After exploring how participants would approach different situations, we asked them to navigate the Inclusive Language Guide. This helped us understand whether the structure matched their expectations when looking for support. 

Several participants were unsure where the practical guidance began. Some mistook the introductory blog post for the guide itself, while others expected the examples and practical advice to appear much earlier on the page. 

Although participants valued the context and principles behind the guidance, they prioritised practical information they could apply immediately. The examples within the guide were consistently highlighted as the most useful content. 

The research continues to inform improvements to the Inclusive Language Guide 

Our findings have allowed us to identify some key areas for improvements to the Inclusive Language Guide.

Making practical guidance easier to find 

One area of focus is the guide’s landing page. We are reviewing its structure so that practical guidance is easier to find while retaining the contextual information that participants found valuable. 

The findings highlighted that the current structure does not always support the way colleagues use the guide. Participants wanted to move quickly to practical guidance while still being able to access the wider context when needed.

We are also reviewing page titles, navigation labels and the ordering of content to make it clearer where the guide begins and how it relates to the wider Editorial Style Guide. We’re considering how to better support both first-time and repeat use, making it easier for people to quickly find practical guidance while still providing the context and principles that many participants found valuable.

We’ll also consider how best to surface practical examples and support the different ways colleagues use the guide, whether they are reading it for the first time or returning to quickly check terminology and writing conventions.

Improving connections with related guidance 

The research also prompted us to consider how the guide connects with related guidance across the University’s web estate. 

Participants frequently expected to find inclusive language guidance within EDI resources or other specialist guidance. Improving discoverability means not only improving the guide itself, but also making sure it is connected to the places colleagues naturally go when seeking support.

As outlined in Emma’s blog, since completing the research, we have worked with colleagues responsible for complementary guidance across the University’s web estate to improve signposting between services.

This work will help ensure colleagues can find appropriate support regardless of where they begin their journey. 

Scenario-based research helped us understand behaviours 

The research reinforced an important principle in content design: guidance needs to be easy to find as well as useful.  

By combining colleagues’ existing experiences with scenario-based activities, we gained a better understanding of how people approach inclusive language decisions. These insights have already informed improvements to the Inclusive Language Guide and will continue to shape how we approach future content research.

QuillMark: The Technical Details Behind My Drupal Style-Guide Assistant

I previously introduced QuillMark, a tool I built to help web publishers check content against the University’s editorial style guide. In this post, I explain the technical design and methodology behind the prototype.

Introduction

My name is Shlok, and this summer I am working as an AI and UX Innovation Intern in the Information Services Group (ISG). During my internship, I have been developing QuillMark, a prototype tool designed to help web publishers check their content against the style guide. This post outlines the system design, methodology and AI pipeline behind the prototype, including how deterministic checks, language models and a MapReduce approach work together to improve reliability while keeping publishers in control.

 

QuillMark highlights violations in your existing text rather than generating an entirely new version. This:

  1. avoids the risk of AI making unnecessary changes to unrelated parts of the page, which would require a manual audit each time
  2. allows publishers to clearly see each violation and accept or reject suggested changes
  3. reduces output tokens, lowering both hallucination risk and cost

 

Problem: The style guide contains a large number of rules, around 60–70. If all rules are sent to the LLM at once, the model may not check the text against every rule reliably.

 

Research: https://openreview.net/pdf?id=R6q67CDBCH

 

Solution → MapReduce (Divide and Conquer: Map → Collapse → Reduce)

 

System Design and Methodology

 

Pipeline:

  1. Deterministic checks to flag blatant violations
  2. Fine-tuned small model flags likely obvious violations
  3. Large model audits the final page deeply for complex deeper violations

 

The deterministic checks are used to flag “obvious” violations that are directly detectable using regex or code. For example, this could include the use of a forbidden word, or a spelling convention such as using “benefited” instead of “benefitted”. These issues can be detected using a direct search of the page and do not require AI.

 

The second stage uses a smaller fine-tuned model. This model would be fine-tuned on the Style Guide rules and training examples. It would again be used to detect relatively “obvious” violations, but ones that are not simple enough to be detected reliably through deterministic checks. These issues are not highly complex, so they should still be visible to a smaller model.

 

After the publisher fixes the deterministic violations and the violations found by the fine-tuned small model, there are likely to be remaining deeper and more complex issues that both previous approaches did not pick up. The final large model can then focus mainly on these more complex issues. This means it spends less attention on obvious violations and more attention on issues that are easier to miss.

 

This improves the process in two ways.

 

First, it allows the final large model to focus on deeper issues. If the whole page was passed directly to the large model from the beginning, it would likely flag many obvious violations first and might miss some subtler issues. By fixing the most obvious issues earlier, the large model can focus more on smaller, judgement-based, or easily missed problems.

 

Second, it may reduce cost compared with running the large model multiple times. One alternative would be to run the large model once, fix the obvious issues + some deeper ones it finds, and then run the large model again to find the remaining deeper issues. In the first run, the large model may spend much of its output on obvious violations. In the second run, after those obvious issues are fixed, it may be able to look more deeply. The proposed approach tries to achieve a similar effect more cheaply by using deterministic checks and a smaller fine-tuned model before the final large-model review.

 

This pipeline is shown in Figure 1.

 

Style Guide Violation Detection Pipeline
Figure 1: Style Guide Violation Detection Pipeline

 

Using MapReduce to reduce instruction-following degradation

Based on personal experience with large language models, as well as existing research such as the Curse of Instructions, we cannot rely on a large language model to follow a very large number of rules at once. In this case, the style guide contains around 60–70 rules. If all of these rules are passed to the model in a single request, the model may not check the page against every rule properly. It may follow some rules, ignore others, or miss violations because there are too many instructions to apply at the same time.

 

To solve this, I propose using a MapReduce approach, which is a divide-and-conquer technique: Map → Collapse → Reduce.

 

The rules would be broken into smaller chunks, with each chunk containing a fixed number of rules. Each chunk would then be sent as a separate request, asking the model to check the page only against that smaller set of rules. This allows the model to focus on fewer rules at a time and check them with more attention.

 

After all the chunks have been processed, the responses would be gathered together and combined into one report. This is the collapse stage. The combined findings would then be passed through another large language model request in the reduce stage. This final request would clean up the output by removing duplicate findings, resolving overlapping issues between chunks, and filtering out possible false positives or hallucinated violations.

 

This approach makes the checking process more reliable because it avoids asking the model to apply all 60–70 rules at once. Instead, each group of rules is checked more carefully, and the final report is produced by combining and cleaning the results. This helps ensure that all rules are checked with more equal rigor, rather than relying on the model to remember and apply every rule in one large prompt.

 

The MapReduce approach is shown in Figure 2.

 

MapReduce Approach for Style-Guide Rule Checking
Figure 2: MapReduce Approach for Style-Guide Rule Checking

 

Update: Large-Model Cleanup from the Reduce stage was excluded from the prototype as the probability of duplicates occurring is low and the cost of such a cleanup is relatively high with minimal benefit.

 

Read more about the technical system plan: https://blogs.ed.ac.uk/website-communications/building-quillmark-testing-a-drupal-style-guide-assistant-with-web-publishers/

Watch a video demo: https://media.ed.ac.uk/media/t/1_yozad8ie

 

Future enhancements:

  • Strengthen Deterministic Regex-matching with more training data (web pages)
  • Audit deterministic findings using a cheap, local AI

Building QuillMark: testing a Drupal style-guide assistant with web publishers

I introduce QuillMark, a Drupal prototype designed to help web publishers check content against the University’s editorial style guide. I reflect on how I developed and tested the tool with publishers, focusing on what their feedback revealed about usability and how it will shape the next version.

Introduction

My name is Shlok, and this summer I am working as an AI and UX Innovation Intern in the Information Services Group (ISG). During my internship, I have been developing QuillMark, a prototype tool designed to help web publishers check their content against the style guide. This blog post explains the problem I was trying to solve, how I built and tested the prototype, what I learned from publishers and what I plan to do next.

 

The problem

Our style guide contains more than 60 rules covering areas such as spelling, punctuation, tone, formatting and terminology. Previous research showed that web publishers found it difficult to remember and apply every rule while writing and editing content.

 

This is understandable. Publishers are often working under time pressure, and checking a page manually against dozens of rules adds a significant cognitive burden. Even experienced publishers can miss small details, particularly when they are concentrating on the meaning and accuracy of the content.

 

I began exploring whether a tool could make this process easier. The aim was not to replace publishers’ judgement or rewrite their content automatically. Instead, the tool would identify possible style-guide violations, explain the relevant rule and allow the publisher to decide whether to apply or dismiss each suggestion.

 

This became QuillMark.

 

Read more about the Style Guide

 

Why I built the prototype in Drupal

Drupal is where publishers already create and edit web content, so it made sense to build the prototype directly into that environment rather than create a separate tool.

Before starting development, I attended Drupal in a Day to understand the platform, its content-editing interface and how a custom tool could fit into an existing publishing workflow.

 

Building QuillMark within Drupal meant publishers could check their content without copying it into another system. It also allowed me to test the tool in a realistic environment, using the same types of fields, buttons and interactions that publishers encounter in their day-to-day work.

 

Read more about Drupal in a Day

 

Planning the checking process

Before writing the prototype, I planned the system in detail.

 

One of the main design decisions was that QuillMark should highlight individual issues in the existing content instead of generating a completely rewritten version. Rewriting an entire page could introduce unnecessary changes and would require the publisher to audit every sentence. Showing individual findings makes it clearer what the tool has identified and keeps the publisher in control. This also prioritises the HITL (Human-in-the-Loop) framework, which I intend to use in my AI-based projects to ensure that people remain involved in reviewing decisions made by AI.

 

The system design proposed three types of checking:

  1. Deterministic checks for clear violations that can be identified using code or regular expressions, such as prohibited words or spelling conventions.
  2. A smaller AI model for relatively straightforward issues that require more context than a simple text search.
  3. A larger language model for more complex, judgement-based rules involving areas such as tone, structure or plain language.

 

The plan also used a MapReduce-style approach. Rather than asking a language model to apply all 60–70 style-guide rules in one prompt, the rules would be divided into smaller groups. Each prompt could then concentrate on a limited number of rules before the results were combined. This was intended to reduce the risk of the model overlooking instructions because it had been given too many at once.

 

The original design included an additional AI step to remove duplicate findings and resolve overlaps. I left this out of the prototype because duplicates were expected to be relatively uncommon, while the extra model request would increase cost and complexity.

 

Read more about the technical system plan: https://blogs.ed.ac.uk/website-communications/quillmark-the-technical-details-behind-my-drupal-style-guide-assistant/

Coding the prototype

I developed the prototype iteratively, beginning with a small number of style-guide rules and expanding the checks once the basic workflow was working. For development, I used Codex. I gave it my plan and asked it to design the architecture. After reviewing it, I asked it to proceed with the implementation.

 

The first stage used deterministic checks for issues that could be found reliably without AI. These checks searched the content for recognisable patterns and returned the location of the issue, the relevant style-guide rule and, where appropriate, a suggested correction.

 

The second stage used a large language model to identify issues that depended more heavily on context. I divided the rules into smaller prompt groups so that the model could focus on a manageable set of instructions during each request.

 

The prompts asked the model to return structured findings rather than a rewritten page. Each finding needed to include enough information for QuillMark to:

  • identify the relevant content;
  • explain the problem;
  • show the related style-guide rule; and
  • suggest a possible correction.

 

Publishers still had to approve or dismiss each finding. This human-in-the-loop approach was intentional. Style rules can depend on context, and an automated suggestion will not always be appropriate. QuillMark was designed as a decision-support tool, not an automatic editor.

 

Watch a video demo: https://media.ed.ac.uk/media/t/1_yozad8ie

 

Designing the usability test

Once the first iteration was complete, I adapted an existing usability-testing template to create a test for QuillMark.

 

I conducted three sessions with web publishers. Participants were asked to work through a series of tasks covering the main parts of the prototype, including running the checks, understanding the two stages, locating an issue in the editor, reviewing a rule and applying or reverting a suggested fix.

 

I observed how participants used the interface, where they hesitated and whether the information on screen matched their expectations. I also asked how they might use the tool as part of their normal publishing process.

 

The purpose was not only to find technical bugs. I wanted to understand whether the concept itself was useful and whether the interface communicated how the tool was intended to work.

 

What publishers told me

The publishers thought the tool would be useful as a final check before publishing. They said they would still use their own judgement and would not automatically accept every suggestion.

 

Most of the feedback was about the usability of the tool. Some parts were unclear, such as the difference between the two stages and some of the button labels. There was also too much information on the screen, and some issues were repeated. The publishers also wanted to see what a suggested fix would change before applying it.

 

What I learned

The testing suggested that the core idea is useful, but the user experience needs to become simpler.

 

Publishers do not necessarily need to understand which checks use regular expressions and which use an AI model. They need to know what issue has been found, why it matters, what the proposed change is and what action they can take.

 

The sessions also reinforced the importance of designing for scanning. More explanation does not always create more clarity. In a publishing workflow, concise labels, clear states and well-grouped findings may be more valuable than displaying every piece of supporting information at once.

 

Next steps

My next step is to refactor and review the prototype code before making changes based on the usability feedback. This will include reviewing the AI-generated code to check whether each section is needed, whether the same result could be achieved more simply, and whether any parts repeat code that already exists. I will compare the generated code with the rest of the codebase so that it follows the same structure and does not add extra functions or files without a clear reason.

 

Where the code is longer than needed, I will simplify it by removing repeated checks, combining similar sections, and reusing existing functions. I will also keep individual files to a manageable size, generally no more than 1,000 to 3,000 lines depending on the file. Larger files will be split where this makes the code easier to follow, and hardcoded rules will be moved into the database where possible.

 

The next iteration will focus on:

  • clarifying or simplifying the two-stage workflow;
  • reducing repeated and overwhelming findings;
  • improving button labels, status messages and colour distinctions;
  • previewing changes before they are applied; and
  • making the interface more concise and easier to scan.

 

Facebook’s Design Didn’t Evolve—It Regressed

14 July 2026 at 12:19
Facebook didn’t get worse overnight—it optimized itself into confusion. What started as the clearest social experience on the web slowly became a noisy, unpredictable system users no longer fully understand. This is what happens when engagement wins over usability—and it’s a warning for every designer building products today.

What I learned at UX Scotland 2026

UX Scotland is an annual two-day conference for people working in user-centred design. This year, I went along to the John McIntyre Centre to hear about the latest developments in UX. In this post, I’ll write about my highlights from the conference.

Object-oriented UX in action: rebuilding a public sector website with structured content

Joey Gartin presenting at UX Scotland 2026. Slide shows text System model in action. Slide has a strange pink and yellow tinge to it.

Joey Gartin presenting at UX Scotland 2026. Weird slide colours courtesy of my phone.

The challenge: dividing up content and naming groupings

Out of all the talks I saw, this was the one that I found the most interesting. Joey Gartin, a content designer at Renfrewshire Council, told a story about a multiyear project to redesign the council’s website. It started with a familiar situation: the council had a large website that people found confusing and hard to use. Users couldn’t find things, and when they could, it wasn’t always easy to understand.

Renfrewshire Council needed to work out a better way to divide up and group the content on their website. Then they needed to name those groups in a way that made sense to people.

I enjoyed hearing about this because this is one of those problems that never seems to go away. When we work with digital content, whether that’s designing a homepage or tidying up a shared set of folders and files, we’re constantly looking for ways to group things. Then we’re also constantly looking for names for those groups that will make sense to people.

OOUX involves favouring nouns over verbs

Renfrewshire Council approached the problem using methods from Object-Oriented User Experience (OOUX). This is an approach to digital product design that focuses on establishing the things (or ‘objects’) that matter to your users. The approach emphasises the importance of working out the nouns involved in your service before you move on to the verbs. I’d read a bit about it a while ago on Duncan Stephen’s blog:

Duncan Stephen writes about how the Scottish Government have used OOUX approaches

The argument in favour of focusing on nouns is that this approach is more closely aligned to how people think. When you walk around a supermarket, you have a shopping list of objects. To reflect this, a supermarket is typically organised into sections that reflect broader groupings of those objects: Fruit, Vegetables, Pasta, Cake, and so on.

In a similar way – the theory goes – people often approach websites with a list of nouns in mind. They are therefore scanning webpages for keywords that match these nouns.

But many council websites use a hybrid of verb-based groupings and noun-based groupings. So you get sections called Pay, Apply, Report and Request forming a core part of their information architecture. Here are two examples.

The City of Edinburgh Council:

Screenshot of the homepage of the Edinburgh Council website, showing links reading Click to pay, Click to report, and Click to request. Other panels show links to Council tax, Bins and recycling, and roads, travel and parking.

West Lothian Council:

Screenshot of the homepage of the West Lothian Council website, showing links reading Pay for it, Apply for it, Report it. and Request it.

 

Renfrewshire have not adopted this model. While their page titles often start with verbs, the groupings for these pages overwhelmingly use nouns:

Screenshot of the homepage of Renfrewshire Council website, showing sections titled Council tax, Bin collection day, School dates and Renfrew Bridge.

Renfrewshire Council

Joey cited the City of Sydney as another public body using this approach:

Screenshot of the homepage of the City of Sydney website, showing sections titled Frequently accessed, Waste and recycling, Building and construction

City of Sydney

OOUX helped Renfrewshire Council to create templates

With your objects figured out, the next step of an OOUX approach is to work out:

  • the relationships between objects
  • the calls to action that objects offer users
  • the attributes that make up objects

Joey talked us through this work. He also talked about how Renfrewshire Council used this as a foundation to create templated designs for common page types.

For example, take these two pages, both categorised as ‘service requests’:

Because they’re both service requests, you can see common sections on both pages.

For example, on the ‘Report a housing repair’ page, you have sections called:

  • Before you report a repair
  • How to report a repair

And on the ‘Graffiti’ page, you have:

  • Before you report it
  • How to report it

Joey showed us how this works behind the scenes. Renfrewshire Council use custom content types in Drupal to help structure the content writing process. So if you’re a content editor writing a service request page, you’ll be prompted to add a section on ‘Before you report it’. This helps content writers know what they need to include in a page. It also brings consistency to the website, which ultimately benefits users.

My reflections on how templates could work at Edinburgh

At Edinburgh, we use content types in our Drupal-based CMS EdWeb to create News and Event pages. So it was interesting to see another public sector website using a more extensive set of content types in Drupal. It made me reflect on how this approach might work for us. The context is so different: at Renfrewshire, it sounded like they had a more centralised approach to content management. The bulk of Edinburgh’s content publishing model is more devolved, which makes templated approaches to content more challenging. But there are clearly benefits available when you can make it work.

So lots to chew on from this talk, and it was a fun presentation to boot.

Other highlights

Sara Wachter-Boettcher on burnout

Sara Wachter-Boettcher’s keynote, “You don’t need more grit: breaking the burnout cycle in UX” was a run through of how and why burnout affects designers working in tech. It was a useful reminder to appreciate the limits of our influence within an organisation, and to be careful not to attach too much of ourselves to our job. It was cool seeing Sara in person. Her book Content Everywhere was one of the first things I read when I started in my role, and I thought it was really good.

Sara Wachter-Boettcher

Craig Abbott on AI and accessibility

Craig Abbott spoke about the risks and opportunities of using large language models for website design and fixing accessibility issues. One point that stood out was that LLMs are trained on some pretty ropey data. WebAIM estimate that 95% of the top million websites have detectable WCAG 2.2 failures, and these are the kinds of site that LLMs are trained on. So if you ask an AI to create a website, it will often produce something that superficially looks ok but is packed with accessibility problems.

The WebAIM Million: The 2026 report on the accessibility of the top 1,000,000 home pages

Craig showed us an example of a website he’d quickly created with AI and talked through the accessibility problems. He also showed us his other experiments. He’d had inconsistent results using AI to detect heading level failures or write alt text that took the context of an image into account. But he’d had some successes with using AI to complete more technical tasks with testable results.

Craig Abbott

James Chudley on digital sustainability

James Chudley presented on bringing sustainability practices into digital design. The highlight for me was seeing his experiments with visualising page weight across a website. Taking inspiration from an infographic using colour bars to illustrate rising global temperatures, James had applied a similar idea to visualising where a website is using more energy-intensive design choices.

James has posted his slides here:

James Chudley: Presenting ‘Beyond human centred design’ at UX Scotland

Another great conference

UX Scotland was a good chance to hear from some leading lights in the field, catch up with people and hear about some case studies that are relevant to us at Edinburgh. The two days were well organised and it gave me a lot to think about as we continue to support UX and content work at the University.

UX Scotland

Information Resources Intern’s first lessons: Time to unlearn “I like the look”?

By: Fan Fei
30 June 2026 at 16:12

My time at the Forrest Hill office as an Information Resources intern began four weeks ago under Edinburgh’s summer brilliance. Working as a part of the Portal Services team, I am responsible for creating and improving existing information resources about MyEd (the University’s online portal), including editing EdWeb2 pages, producing instructional videos, and refurnishing help and support SharePoints.  

Experiences designing promotional graphics, videos and syllabi for societies and events since high school made me a veteran of CapCut and Canva. It didn’t take me too long to get the gist of less familiar platforms like SharePoint and EdWeb2 (the University’s content management system) either, thanks to the self-paced online training courses. Freshly equipped with all the essential tools, I hit the ground running with high spirits. Like in any independent projects and commissions I did before, my creative visions, personal judgements and the mere intuition of “I like the look” carried me through the planning and proposal stage. I audited the existing EdWeb2 pages about MyEd to propose improvements, planned the refurbishment and new structure of the SharePoints, and storyboarded for the upgrade of the “How to use MyEd” video. Everything progressed beautifully smoothly and swiftly. The big frameworks and ideas were set up and ready to go.  

 

When “I like the look” starts to fail  

But as they say, the devil’s in the details; mine was waiting right around the corner at the execution stage.  

I was the most ambitious and eager with the upgrade of the “How to use MyEd” video, as it would be my first time creating official video for an organisation rather than for smaller commercial causes. My initial aspiration for upgrading this video was to make it as eye-catching and engaging as the promotional videos made by big companies like Substack: intricate transitions, animations, and rich colours. I wanted it to not only be informative, but charming enough to excite the new students about university life and this new portal system they’re starting it with. I spent nearly a week adapting to the PC version of CapCut, learning new transition techniques, and adjusting the properties of millions of key frames for about the millionth time until the movements are natural and the timings are perfect. With fondness and personal satisfaction towards the visual, I submitted the first draft for review.  

My manager responded:  

“Have you reviewed it with the accessibility checklist?” 

That’s when the design I was so proud of began falling apart like a house of cards. It became clear to me the first time that I’m no longer designing resources for a class of 30, but a community of 53,000+ students within a large institution. “I like the look” hardly stood valid against objective, rational standards like the WCAG (Web Content Accessibility Guidelines) and the University’s editorial guide for effective digital content. The sliding animations I designed turned out to be rather disruptive, as they create moving backgrounds under the text. This compromised the minimal contrast requirement (as text might blur into background) and created more challenges for users with attention or vestibular disorders. After re-reading the WCAG sections on video accessibility, I reduced the speed and frequency of motions on screen, avoided all-caps titles (which troubles users with dyslexia), and carefully checked font size and color contrast section by section.  

Around the same time, I came across Katie’s wonderfully illuminating blog on when videos and images genuinely add value to digital content: 

Videos and images – when do they add value and not just page weight? – Website and Communications Blog 

Although the blog was more about websites, its message inspired me to remove stock footage in my videos, which I previously employed for pure aesthetic reasons to make sections appear less static. Not only did they not add much informational value, but they also increased the video file size and created more moving backgrounds that hinder accessibility.  

Another lesson came through collaboration. Unlike society and event projects where I made every creative decision myself, this video formed part of the Pre-arrivals team’s “How to” series and therefore needed to follow an established visual identity. I will need to adapt the current draft to their prescribed templates, formats and requirements. With the same logic of branding in mind, I also replaced all colors in the video with colors from the University’s official branding palette. This way, in the long term, it could seamlessly fit into university websites, whether uploaded to Media Hopper or embedded into EdWeb pages.  

It would be dishonest to say it wasn’t a little heartbreaking to undo and cut work I had spent hours creating. Nonetheless, I came to understand that information resources differ from creative projects exactly because clarity, effectiveness, and accessibility must always come before visual aesthetics. I began seeing various guidelines and requirements less as constraints. They taught me a more rational and methodical approach to creating information resources than relying solely on the easy intuition and judgement of “I like the look”. Where I designed with my own palate and assumptions before, now I truly began my journey of designing for others. 

 

But does “I like the look” truly have no merit at all?  

Surprisingly, I would argue that it still does. “I like the look” is not something to unlearn completely, but it should be treated as the beginning rather than the end of the design process. From a different perspective, I have come to think that that instinctive thought is where UX (user experience) thinking begins. 

When I first opened the Notifications Service SharePoint that I was supposed to redesign, my immediate reaction was “I don’t like the look.” Except this time, rather than concluding the thought and jumping straight to redesigning based on personal preference, I attempted to find a rational “Why” by referring to all the guidelines and effective digital content tips I have learnt.  

The SharePoint was comprehensive, containing all the information staff would need to understand what notifications are and how to send different types of notifications. Yet as a first-time user that knows near nothing about the Notifications Service, I found it difficult to know where to begin. The original homepage was filled with large blocks of text and links with equal visual weight regardless of importance; Navigation of all the pages and materials relied heavily on users have prior knowledge on terminologies like “Notifications Template” and “Emergency Notifications”, and there was little visual hierarchy to guide attention or suggest the next step.  

That’s when I realised when users say they “like the look” of a website, they are often responding to not just colours or typography, but how intuitive, organised and effortless it feels to use. I decided not to simply edit the existing pages, but to rebuild the site’s structure in a separate testing environment.  

For instance, instead of letting the homepage act as a page of plain text, I plan to make it a clearer starting point for all types of users. I intend to include an image of Notifications Backbone (the system used to send notifications) alongside a direct button to access it. This way, experienced staff can access the system immediately, while first-time users are given some introductory visual and textual context before they continue exploring the site. The content changes minimally; my focus is to reorganise them around users’ journeys. Similarly, rather than presenting users with a long list of pages whose titles assume prior knowledge of relevant terminologies, I plan to reorganise the content into a small number of category cards based on what potential tasks users are trying to accomplish. I also hope to introduce a clearer hierarchy within the navigation menu, so that related pages are grouped more logically, and any information can be found more intuitively.   

Original homepage  v.s. New draft hompage 

  

Original information list  v.s. New draft categorisation  

New draft menu bar  

Being a student unexpectedly became an advantage in this process. Since I had similarly little prior knowledge of the Notifications Service as a new user, every point where I became confused while visiting the original SharePoint highlighted somewhere a common struggle might occur. Looking at the site with fresh eyes helped me identify problems that familiarity can sometimes hide. 

 

From “I like it” towards “We like it”  

But of course, just as I learnt with the review of the MyEd video draft, the work does not end once I think the design “looks right”. Everything I design  – SharePoint, websites, and videos – still need to be checked against the accessibility standards, follow editorial and branding guidelines, and most importantly, be tested with the people it is actually designed for (e.g. test MyEd video with students, the Notifications SharePoint with staff). Nick’s excellent guide on running usability tests has been particularly helpful in shaping how I plan to continue with the next steps.  

How to run a usability test – Website and Communications Blog 

“I like/dislike the look” has carried me through countless creative projects before. Four weeks later, I don’t think that instinct has been unlearnt but simply found its proper place. Rather than a justification for me to blindly follow personal preference and biased judgements, it has become the starting point for me to ask better questions about accessibility, usability, and the needs of the people I’m designing for. That is the perspective I’ll continue carrying into MyEd information resource that I will continue to create and develop throughout the rest of the internship.  

 

 

Moving Beyond UX: The Rise of the Agentic Experience (AX) Designer

23 June 2026 at 12:41
As AI agents quietly take over workflows, a new discipline is emerging: AX (Agentic Experience) Design. The designers who thrive in the next decade won't just craft experiences for humans—they'll define the rules, guardrails, and invisible systems that autonomous machines use to make decisions.

How to run a usability test

In our May Content Improvement Club session, we focused on how to run a usability test. We ran through the basics of putting a script together, watched a clip from a test and had a go at prioritising some issues.

Two problems for people working with content

We started this session by outlining two problems for people working with content.

We don’t see people using our websites

The first problem is that we don’t typically see people using the things we create. We build a website based on our best assumptions of what our users need and how we think they will interact with it. Then at some point in the future, someone uses the site. They might find it easy to use. They might find it difficult. But we don’t know, because we don’t get to see.

It’s hard to self-assess our own website

The second problem is that it’s hard to self-assess a website that we’re already familiar with. When we look at the site, we bring our understanding of the structure and the context it sits in. We know what the acronyms mean. We know where to find the contact form. We know where the links go.

This isn’t necessarily the case for a user coming to the site for the first time.

Here’s Steve Krug, author of Rocket Surgery Made Easy:

“If you’re building something, you’re not going to be able to see where it’s going to confuse people. It’s not going to confuse you. You know too much about it.”​

Steve Krug interviewed on the Brave UX podcast

This is sometimes known as the ‘curse of knowledge’. Erika Hall, author of Just Enough Research, explains:

Whenever we learn things, we forget what it’s like not to know those things. […] The more you know, the more you expect other people to know. And if you become a real expert in a topic, forget it.

Erika Hall writing in the Mule Design Studio newsletter: Don’t chicken out about talking to people

Usability testing is a way of addressing these problems. It gives us a chance to see what it’s like for someone visiting our site for the first time and it counteracts the curse of knowledge. That gives us valuable insight into where a design is working and where it isn’t.

A quick definition of usability

Before we go any further, a quick word about what we mean by ‘usability’.

In short, usability refers to how easy, effective and satisfying something is to use.​

The cover of Don Norman’s book ‘The Design of Everyday Things’ features a coffeepot with a spout and handle on the same side. This would score poorly in a usability test as it wouldn’t be easy, effective or satisfying to use.

The cover of the Design of Everyday Things, featuring a coffeepot with a spout on the same side as the handle.

One of French artist Jacques Carelman’s impossible objects, as featured on the cover of the Design of Everyday Things.

ISO definition

There’s also an international standard (ISO 9241-11) definition of usability:​

“The extent to which a product can be used by specified users to achieve specified goals with effectiveness, efficiency and satisfaction in a specified context of use.”​

ISO definition of usability

This is a more technical definition, but it’s worth bearing in mind. It emphasises the fact that we need to think about specified users and their goals. That will come up when we get to writing our script.

A simple usability test

The simplest usability test you can do takes about 5 minutes:

  1. Grab a colleague / friend / family member who doesn’t know your site​.
  2. Sit them at a computer and open the homepage of your site.
  3. Set them a task. For example: “Imagine you’re a student and you want to get a replacement student card”.​
  4. Sit quietly and watch as they try to complete the task.

This doesn’t take much effort, and sometimes you uncover a usability issue or two.

In Content Improvement Club, we looked at how to run a more extensive usability test. But we wanted attendees to know from the outset that usability testing doesn’t have to be a big complicated process. It can still be beneficial even if you do it on a small scale.

There are three roles in every test

In a usability test, there are three roles:

  • A moderator, who reads the script, sets tasks for the participant and asks questions​.
  • A participant​, who completes tasks on a website​, thinking out loud​ as they do so.
  • An observer​, who watches and takes notes​.

There can be any number of observers, who watch the test live or on a recording.

We watched an example of a usability test

The best way to understand how a usability test works is to watch a recording of one.

Here’s Steve Krug demonstrating how he runs a test:

Usability Test Demo by Steve Krug (YouTube video, 24 minutes)

In Content Improvement Club, we watched a short clip of someone completing a task on a library website at a UK university.

This was the task:

You’re working on a presentation with four other people from your course. You need to find a study room in the library for the four of you this weekend. Can you find out if a suitable study room is available at the library?

We noted down the issues that we saw, and then we shared them on a Microsoft Whiteboard. Even though it was a single task, there were a lot of issues, which is fairly typical. That’s why it’s often helpful to follow your observations with a prioritisation exercise.

We prioritised the issues

In the session, we demonstrated how you can prioritise issues using a matrix like this:

A matrix showing Easy to fix and Hard to fix on the vertical axis, and minor issue for users and major issue for users on the horizontal axis. Blank sticky notes are in some quadrants.

To place things on the matrix, you make a quick assessment of how significant the issue is. Then you assess how easy it would be to fix. The issues you typically want to focus on first are those in the top right quadrant: major issues that are easy to fix.

It can be helpful to have a set of questions to establish whether an issue is major or minor.

These are adapted from Dave Travis’s work on Red Route usability testing:

  • Does the problem occur in a task that is highly important to you or your users?
  • Is the problem difficult for users to overcome?​
  • Did multiple participants experience the same problem?

Red route usability: The key user journeys with your web site (Dave Travis / archive.org)

Using questions like these give you a shared set of criteria when assessing the significance of usability issues in a group.

But we’re getting ahead of ourselves. How did we get to this point?

How to write a script

Before you can run a series of tests, you need a script.

Use a template

We shared our usability testing script template, which draws on the work of Steve Krug:

Research brief and usability testing script (DOCX, 51KB)

The template includes:

  • a brief, where you articulate the goals of the testing
  • a lead-in script, which you read to participants before testing begins. This covers what the session will involve and sets expectations. ​
  • a set of tasks, presented in a table with two columns. One column contains the task, and another the expected path to complete the task.​

Come up with tasks

We start by identifying tasks and then develop them by adding a scenario.

We practised this in the session. Attendees suggested tasks that someone might need to carry out on a university library website. For example:

  • Check opening times
  • Find a book on your reading list

Develop tasks by adding scenarios

Next, attendees developed these tasks by adding a scenario.

For example, “Check opening times” becomes:

  • You’re a student and you want to check the opening times of the library during the winter break. How would you go about doing this?

“Find a book on your reading list” becomes:

  • You’re a new student and you want to find the book “Campbell’s Biology”, which is on the reading list for your course. Can you show me how you would find out if the library has a copy of this book?

Tips for writing good tasks and scenarios

We shared some tips for turning tasks into questions​:

  • Aim for 5 to 10 scenarios per session.​
  • Keep tasks focused: one goal per task.​
  • Frame tasks as realistic scenarios, not instructions.​
  • Add light context to make it realistic.​

We recommend running a pilot usability test before going ahead with a round of testing. This helps you see how the full usability test flows from start to finish. ​It also gives you a chance to refine the script and ​ensures the session runs smoothly for participants on the day.​

Logistics

Ahead of the Content Improvement Club session, we asked attendees if they wanted us to cover any topics in particular. Most of these came under the wider topic of logistics.

When should you test in a design process?​

Test earlier rather than later. Testing earlier in a design process is ideal because it’s easier to make changes based on what you learn. If you’re creating something new, this usually involves creating rough drafts or prototypes, and then later, a high-fidelity version.

Changing a prototype is easy because no one is very attached to it. But when a design is a later stage of development, it tends to be more painful to learn that something isn’t working. By this point, you and your colleagues have usually invested a significant amount of time in an idea. If the fix involves changing something fundamental to the product, you might have to undo work that’s already been done.​

How long does a testing session take?

It varies, but we usually run testing sessions of about 30 minutes.

How many tasks do you set?

In 30 minutes, we can usually fit in between 5 and 10 tasks.

How many participants do you need?

Three to five participants is a good target. Jakob Nielsen argued that five participants is enough for one round of tests. This is because when you run the same tests with multiple people, you tend to see the same usability issues reoccurring. It’s a case of diminishing returns: by test number six, you aren’t usually learning anything new.

Why you only need to test with five users (Nielsen Norman Group)

In Rocket Surgery Made Easy, Steve Krug recommends three participants per round of testing. This way you can run more rounds of tests – for example, by testing an initial idea and then a later iteration of the same design.

How do you recruit participants?

Mailing lists, Teams channels and surveys are great places to put call outs. We advise that you avoid revealing the testing topic or content in advance.​

Aim for participants who reflect real users where possible. This can prove difficult, so don’t let it stop you if you can’t find participants who don’t match the profile of your users.

How do you take notes?

Obviously, you can scribble notes anywhere you like. When we’re running a series of tests, we often take notes in a spreadsheet. This allows us to quickly spot which tasks caused problems for multiple participants.

Usability test results spreadsheet (XSLX, 36KB)

How can we make testing more accessible and inclusive?

Accessibility and inclusivity should be considered from the very start of the testing process. One important step is asking participants early on whether they use any assistive technology, so you can understand their setup and make any necessary arrangements. Our own introduction questions and templates include this for that reason. In the template, this is an introductory question, but if you’re recruiting via a survey, you could mention this there.​

​Having a diverse participant group will usually lead to more useful and representative findings. For example, if you’re testing a student-facing service, including both undergraduate and postgraduate students can help capture a wider range of experiences and needs. Again, this can sometimes be a challenge, but if possible something to aim for.​

Ideally, usability testing should include participants who regularly use assistive technologies. This provides valuable insight into accessibility barriers that might otherwise be missed. However, this can sometimes be challenging due to recruitment costs, specialist panels, or limited budgets.

Some resources that can help:

Acting on the findings​

So you’ve run a round of testing. What happens next?

Involve senior colleagues in playbacks

In 2015, Caroline Jarrett wrote about a phenomenon she and Steve Krug had noticed when working on websites. They would run a set of tests on a website and uncover some usability problems. But then six months later, the problems were still there: no one had gone in and fixed them. To investigate why this was happening, they ran a survey of UX professionals.

Caroline Jarrett and Steve Krug’s analysis of why usability problems go unfixed

Out of 131 responses, the most common reason was that findings from usability testing conflicted with a decision maker’s opinion.

One way to address this is to get decision makers in the room (or on the Teams call) when you watch a highlights reel of test recordings. There really is no substitute for seeing user behaviour first hand. Reading about it in a report doesn’t carry the same weight.

Use the momentum created by testing

Testing creates a shared motivation to fix things​. Use this to your advantage. If you can, take action to fix usability problems while people’s memories are fresh​.

Make small changes first

Sometimes usability testing highlights small problems that are easily fixed. A link that doesn’t go where someone expects it to. An item missing from an A-Z.

Sort these things out first. Small wins like this give you the motivation to sort out the knottier problems.

If you want to find out more

This was a quick introduction to running your own tests. If you want to learn more,  Steve Krug’s introductory guide is a great place to start:

Rocket Surgery Made Easy by Steve Krug (listing on DiscoverEd)

We post about case studies of usability testing at the University on this blog:

The Prospective Student Web Team do the same:

Acknowledgements

Various points made in this blog post are taken from this 2015 post by Neil Allison:

Making usability testing agile

How to hear about future sessions

We promote these sessions via our mailing list. If you’re interested, please sign up:

Join the UX and Content Design mailing list (University login required)

Suggest a topic for a future session

We picked usability testing following a suggestion from the community. We’re keen to continue covering topics that colleagues across the University would find useful. It would be really helpful if you could let us know any ideas you have using this form:

Suggest a topic for Content Improvement Club (University login required)

Other training that we offer

More training is listed on the User Experience Service website:

Training | User Experience Service

Stop chasing, keep researching: Why continuous contextual learning is the only way to build useful AI features

AI development keeps evolving as do the ways people seek to use AI. Traditional software development runs the risk of trying to perfect AI features people won’t use. Revisiting our previous AI research helped me tease out new opportunity spaces for AI features to help with content design tasks.

Last summer, Mostafa Ebid joined the UX team for a summer internship and built an AI assistant tool which integrated the University’s main AI provider ELM into EdWeb2, our Drupal content management system. The idea for the tool came from hearing University describe the difficulties they experienced when publishing web content. The tool included the capability to write content, design content and proofread, and was designed to work by typing prompts in a chatbot interface, right-aligned to the main part of the editorial interface. When prompted, an orchestration of AI agents were triggered to read textual content and use ELM to make suggestions for improvement based on what the publisher had asked for. Improved text was displayed in the sidebar chatbot interface for the publisher to review and consider using.

Read Mostafa’s blog post about how he developed the tool:

Integrating ELM with EdWeb – Building an AI tool for publishers

Initial tests of the tool with publishers revealed some potential, but identified the need to do further tests, specifically to understand limitations around the user interface display and to learn if the tool was useful to publishers in the context of content they were familiar with (as opposed to generic stock content that was used in the first round of tests).

Read my blog post about the initial tests of the tool:

Initial insights from UX testing our Drupal AI content assistant tool 

Mel Batcharj accessibility tested the tool, focusing on keyboard navigability, and Nick Daniels ran tests with two publishers in October last year. I recently reviewed the findings of this research to consider advancements in Drupal AI in the past 12 months, to assess whether the original premise for the tool was still valid, and to think about residual content design challenges AI could potentially help with.

Advancements in Drupal AI have resulted in improved AI features

The Drupal AI Initiative began in April 2025 with a group of participating organisations making a commitment to collaborate to build the future of AI in Drupal. As a result of the initiative, various workstreams began to shape Drupal’s infrastructure to support AI, to experiment with new innovations and to improve the UX.

Read more about this workstream:

Drupal AI Initiative project page on Drupal.org

A more accessible AI chatbot is now available as a Drupal recipe

Our original AI content assistant tool was built into a panel of the editorial interface as a custom build, which worked well when using a mouse, but which accessibility testing showed was restrictive when navigating using a keyboard. Since the tool was built, however, accelerated development in the wider Drupal AI community prompted the creation of an open-source AI chatbot freely available to apply. Adopting this chatbot was preferable as it avoided the need to maintain custom code and it was possible to use it with a keyboard only.

There’s an active Drupal issue to address the need for the AI chatbot interface to be expandable

Several of the participants who took part in Mostafa’s tests of the tool last year found it awkward to scroll through the AI chatbot output as it was presented in the narrow left-aligned interface. The same problem had been noted in the wider Drupal community, and therefore I raised an issue to have this rectified, which is being worked on as part of the AI Initiative task backlog.

We did more tests on our AI tool – this time using participants’ own content

In the first round of tests, four participants were all presented with the same piece of content (on the topic of safety procedures) and asked to use the AI tool to improve it in specific ways (such as writing it for user needs, making link text better, and so on). This approach turned out to be limited, as since participants were unfamiliar with the content, they were unable to assess whether the outputs of the AI tool were an improvement on the original content or not.

In a subsequent round of tests we therefore adopted a looser approach – asking participants to supply a piece of content they were already working on, and then asking them to use the AI tool to improve it to suit their needs. The results from these tests were more indicative of how useful publishers found the tool to make content improvements.

Results of these tests highlighted how the AI tool needed to change

Since we placed participants in a situation where they were using AI on content they knew well and could critique and appraise authentically, the results of these second-round tests gave clearer indications of what worked with the existing tool, what didn’t and where AI could be best applied to help with content design tasks.

Participants didn’t notice the content options in the AI tool, or the help text

The tool contained three different content options: Design Content, Write Content and Proofread, presented in a dropdown menu in its interface.

Initially, participants didn’t notice these options and used the default Write Content option.  When they later experimented with the Proofread option they found no discernible difference between these options in terms of outputs, which led them to believe that a simpler version with a single conversational interaction option would be preferable.

The tool defaulted to reading the content on the page, and working on this when prompted. Participants were initially unclear that this is how it worked, and they didn’t notice the help text to enable or disable this mechanism presented in the tool interface. Taken together, this feedback suggested that a simpler version of the AI chatbot, such as the one from the Drupal recipe, would be easier for publishers to use.

Close-up screenshot showing detail on the AI assistant tool chatbox

Close-up screenshot showing the content options and the help text on the AI tool

The AI tool had some value as a writing partner to suggest restructures to textual content

Responding to the task to experiment with the AI tool to improve their content, the participants quickly got used to how the tool worked, and recognised its use as a writing partner to prompt about their content and receive suggestions for improvement in a conversational way.

Prompts they used to improve a page of content in the body text field included:

  • ‘Rephrase copy to condense, highlight key messages and make it accessible to pet owners looking to join practice’
  • ‘Proofread copy so that it appeals to pet owners non clinicians’
  • ‘Turn this page into web ready content. It needs to be concise, easy to scan, readable to a wide range of audiences’

These prompts resulted in edited versions of the page content, typically including structural elements like headings, bullet points and calls to action, delivered in the chatbot interface.

Screenshot showing the output from the AI tool before and after a prompt to rephrase the content for a specific audience (before on the left, after on the right)

Side-by-side screenshots showing the prompt entered in AI tool to improve content for an audience (on the left) and after (on the right), with the output response to the prompt.

The tool lacked capacity to tweak or iterate on previous versions of content – which participants wanted

Once they had reviewed the tool’s initial outputs, both participants entered conversational turns with the tool, asking it to perform successive tasks on the content it had previously produced, to edit it further, in line with their specific requirements and rules.

Prompts they used to tweak initial AI outputs included:

  • ‘Remove adjectives and exclamation marks’
  • ‘Remove brackets’
  • ‘Remove any unnecessary words, fix typing errors, suggest improvements for SEO’

With every new prompt in the conversation, the tool produced a fresh output, meaning the publisher was left to review a succession of different versions of the edited content, presented in the chatbot interface, without any indication of what had been changed. Participants found it difficult to review edits, as they would usually do when working on a piece of content, in order to compare the ‘before’ with the ‘after’ – ultimately to assess whether the AI tool outputs were to their satisfaction.

They said they would have liked the tool to have presented the edits in a ‘tracked changes’ format that they were familiar with from word processing programmes.

Screenshots showing outputs of the tool before and after a prompt to iterate on improved content

Side-by-side screenshots showing a prompt entered to improve existing content (on the left) and (on the right) the output from this prompt – showing a new version of the content

Participants didn’t really need the tool in the interface as they tended to edit text content elsewhere

When describing their usual content process, participants said they would usually prepare their content in a word processing programme like Microsoft Word rather than edit directly in the Drupal editorial interface. There were several reasons they chose this method – a key ­­reason being the need to involve others to check (and in some cases, sign off) their content in preparation for the website. They were more familiar with referring others to check their content or proofread it when it was in the Microsoft suite, than when it was in the Drupal editorial interface.

Furthermore, Microsoft Word accommodated the addition of comments and tracked iterative changes to pieces of content which was not possible within the Drupal editorial interface. This content preparation habit suggested that while the AI tool was useful to suggest content edits, this would have been more useful before the content was in the interface, and therefore could be achieved by pasting content to be edited into a browser-based AI tool or app (such as ELM).

Within the interface, the tool only had use as a ‘final check’ mechanism, to catch any typos, errors or style misalignments before the content was ultimately published.

The tool needed to be able handle more than text as pages were typically made of multiple elements

Reviewing the test set-up and comparing it to their usual ways of working with EdWeb2, the participants said the pages they worked on would usually be made up of more than just textual content in the body text field. They would typically work on pages with multi-column layouts, making more extensive use of Drupal paragraphs or including structural elements like accordions, feature boxes and cards. They were interested to know how the tool may make appraisals or suggest improvements for those sorts of pages to help them arrange their content in appropriate ways.

We identified new opportunities for AI content publishing features

Extrapolating on the feedback from the tests, several use cases and scenarios for applying AI to content design tasks emerged, which will help inform our ongoing work to apply AI to make content design tasks easier for publishers.

AI-assisted content structuring

Describing their typical content writing workflow, participants said they found it difficult to move from a text-based editor like Microsoft Word into the Drupal editorial interface where they needed to make use of Drupal Paragraphs as well as page elements like accordions to structure the content. Potential areas for AI development could therefore include:

  • A mechanism to convert textual content into appropriate structural elements
  • A way to make suggestions for accordion labels or structure
  • A feature to ensure uniform creation of cards or feature boxes
  • A means of cross-checking style consistency of pages made of multiple elements or with a specific layout

AI- assisted content design for SEO/GEO/AEO

Having their content picked up by search engines or being machine read was something participants wanted, but they were unsure how to write, tag and structure their content to make this happen effectively. Potential areas for AI development could therefore include:

  • A way to have their content analysed for SEO effectiveness, based on signals like content quality, scannability and key word alignment
  • A mechanism to suggest content changes aligned to specifically defined SEO goals and target user engagement measures

AI assisted content reviews – against specific style conventions and contexts

As well as ensuring their content followed the rules of the University’s Editorial Style Guide, both participants mentioned other conventions that they needed to apply to their content, that existed at a more local website level. For example, one participant’s site had a rule not to use brackets or exclamation marks, or to overuse adjectives, so they would have found it helpful to have a way to cross check content against these rules before publishing.

The Drupal Context Control Center is an emerging feature designed to handle the application of context rules within a site across various scopes and use cases, and Drupal AI Content Review is a related feature, designed to appraise content against given context rules and conventions. Together, these Drupal features may be a good fit to help University web publishers make use of AI to shape their content with the uniformity they require.

Read more about the Drupal Context Control Center and its development in my recent blog post:

Think like a machine: How building a Drupal context-handling feature is providing a new lens of content design and style rules

Read about AI Content Review on Drupal.org

We’re changing the direction of the AI tool based on what we’ve learned

Last summer it seemed certain that an in-interface AI chatbot content assistant helper was what we needed to build. A few improvements to the UI to make it expandable and navigable with a keyboard and it would be ready to go. As it turned out, things had moved on, and these problems were addressed by the wider Drupal community. This meant we could go back to our research findings to re-examine how AI could be best applied to aid content design tasks, and to consider how it could best fit into existing workflows of our publishers to assist them with difficulties they faced. As we continue with internships this summer, we’re excited to re-focus and plan more research to explore some of these emergent opportunity areas. Aligning with in-progress Drupal AI developments, we’re open to learning how we can apply and adopt the work of the Drupal community to our University digital publishing context.

❌
❌