9. Data sources and quality

Data source and collection

The data for this release come from HM Revenue and Customs’ (HMRC’s) Pay As You Earn (PAYE) Real Time Information (RTI) system. They cover the whole population rather than a sample of people or companies, and they will allow for more detailed estimates of the population. More information on the quality of the data and the steps we take to quality assure it can be found in our Quality assurance of administrative data used in earnings and employment from PAYE RTI methodology.

Our statistical practice is regulated by the Office for Statistics Regulation (OSR). The OSR sets the standards of trustworthiness, quality and value in the Code of Practice for Statistics that all producers of official statistics should adhere to. You are welcome to contact us directly with any comments about how we meet these standards by emailing rtistatistics.enquiries@hmrc.gov.uk. Alternatively, you can contact the OSR by emailing regulation@statistics.gov.uk or via the OSR website.

HM Revenue and Customs logo

Coverage

This publication covers employees payrolled by employers only. It does not cover self-employment income or income from other sources such as pensions, property rental and investments. Where individuals have multiple sources of income, only income from employers is included.

The figures in this release are for the period July 2014 to June 2026 and are seasonally adjusted.

Upcoming changes

Following the UK’s withdrawal from the EU, a replacement to the Eurostat geographical classification NUTS regions has been created. The UK-managed classification of International Territorial Levels (ITLs) will replace the NUTS classification in future publications.

Please contact us at labour.market@ons.gov.uk and rtistatistics.enquiries@hmrc.gov.uk if you would like to offer feedback on how the contents can be improved in the future.

Methodology

Our accompanying article contains more information on the calendarisation and imputation methodologies used in this bulletin, alongside comparisons with other earnings and employment statistics and possible quality improvements in the future.

Pre-release data

HMRC grants pre-release access to official statistics publications. As this is a joint release, and in accordance with the HMRC policy, pre-release access has been granted to a number of people to enable the preparation of statistical publications and ministerial briefing. Further details, including a list of those granted access to official statistics by HMRC, can be found on HMRC’s website.

Accredited official statistics

The Office for Statistics Regulation independently reviewed these accredited official statistics in July 2025. They comply with the standards of trustworthiness, quality, and value in the Code of Practice for Statistics and should be labelled “accredited official statistics”.

This is a joint release between HMRC and the Office for National Statistics (ONS).

Strengths of the data

As PAYE RTI data cover the whole population, rather than a sample of people or companies, we are able to use these to produce estimates for geographic areas and other more detailed breakdowns of the population. The methods for producing such breakdowns are under development and we expect to include further statistics in a future release. These statistics can help inform decision-making across the country. They also have the potential to provide more timely estimates than existing measures.

These statistics also have the potential to replace some of those based on surveys, which could reduce the burden on businesses needing to fill in statistical surveys.

Industry Sector Classifications

The industrial sectors in this bulletin are based on the UK Standard Industrial Classification (SIC) codes, as defined by the ONS. These codes have been determined from both the most recent Inter-Departmental Business Register (IDBR) and data from Companies House for each PAYE enterprise.

Large enterprises that cover multiple SIC codes are classified into a single SIC code based on the relative number of employees in each SIC code. Changes to the proportion of employees across SIC codes in large enterprises can result in the enterprise being reclassified to a different SIC code. To obtain the SIC code, we link to the most recent quarterly versions of the IDBR. Once a year when we refresh data for the whole series, the IDBR link is refreshed using the most recent version available, and any reclassifications are then used for the entirety of the time series until the next year.

This means that sector level time series represent the current employers classified in each sector and are less likely to be distorted by employers being reclassified at the enterprise level because of small changes at the lower unit level. However, it also means that these time series may be revised between publications and, in the historical sections of the time series, employers are classified in sectors in which they were not classified at that point in time. However, this method should minimise discrepancies in the data caused by reclassifications and should more easily allow the tracking of job movements between sectors.

Imputation and revisions

RTI data used in this release are extracted in the weeks following the end of the latest reference month. For some individuals, this means payments relating to work done in recent reference months are yet to be received. Rather than wait until all payment returns have been received, we produce timelier measures by imputing the values for missing returns.

For the latest reference month, around 15% of the data are imputed. We refer to this as the “flash” or “early” estimate in the bulletin, because this figure is the most subject to revision as payment returns are received and the imputed payments replaced with actual data.

From our July 2022 publication, two changes were made to the imputation model. A seasonal factor was incorporated into the imputation model. The model was also made more responsive to recent changes to the labour market that would affect the likelihood of a payment existing. The latter change in particular should reduce the scale of revisions seen to the “flash” estimate, but cannot eliminate revisions completely.

Earlier months also contain some imputed data. Some payment frequencies mean that we have not received the relevant payment data more than a month after the reference period. Also, in some circumstances, returns might be submitted late. Therefore, earlier months are also subject to revision, but these revisions are likely to be much smaller because the level of imputation is smaller. The proportion of imputed data for a reference month two months before data extraction is around 1% to 2% of the data.

For the majority of months, post-flash revisions will occur in small amounts gradually each month as more submissions are received. However, all RTI submissions must be received before the end of the tax year. Therefore, for months close to the end of the tax year, these submissions and associated minor revisions that would have accumulated through the year instead need to be received all at once in the final submissions of the tax year. The months of January and February will be most affected by this and see sharper non-flash revisions at the end of the tax year if the imputed submissions are not received by that point. From July 2022, changes were incorporated into the imputation model to try to control for these seasonal differences, as well as other seasonal factors that might affect whether submissions are received through different points of the year. Further information on the impact of the changes to the imputation model can be found in our Impact of imputation changes in employment statistics from Pay As You Earn Real Time Information methodology.

The seasonal adjustment model will also update each month as the model is refined on the latest data available. These adjustments will appear as revisions in the seasonally adjusted data, and in the supporting seasonally adjusted revisions triangle.

Starting with the December 2020 publication, we introduced a new revisions policy. For each publication, we incorporate new input data only for the current tax year and the previous tax year. Revisions to estimates can potentially be made for up to the last two years as data can continue to be received, though updates to data outside of the most recent tax year are minimal.

Changes to the seasonally adjusted data also occur earlier than this limit, as the seasonal adjustment model is refined. The benefit of introducing this revisions policy is that we can use the processing time saved to produce and publish more detailed breakdowns. We capture any new input data referencing earlier years by incorporating data for the whole time series once a year.

Seasonal adjustment

The seasonal adjustment applied in this bulletin follows established best practice. This approach assumes that any seasonal patterns remain broadly consistent over time. If the seasonal pattern changes in strength, this will be represented as greater volatility in the seasonally adjusted figures. Both the seasonal and non-seasonally adjusted datasets are released alongside this bulletin.

Differences compared with other labour market statistics

The Labour Force Survey (LFS) is our survey of households, while Workforce Jobs (WFJ) is based mainly on business surveys for employee jobs, with the LFS covering self-employed jobs. HMRC PAYE RTI data are derived from administrative tax records and only cover payrolled employees.

Each of these three sources are collected and processed in different ways, so we do expect differences in levels (for example, jobs versus people, differing reference periods). Divergence across indicators for more than one period is not unusual. For further information, please see Section 3: Trends and considerations around comparisons in our Labour market overview bulletin.

Back to table of contents



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *