Releasing an AI Model
Preparing and releasing a trained AI model on Hugging Face or a similar hub
Releasing source code instead?
This page covers releasing a trained AI model (its weights). For source code, see
Releasing Open Source. If you are consuming an external open source model rather than
publishing one, see AI Model Licenses.
Releasing an AI model follows most of the same steps as releasing source code. You obtain
organizational approval, confirm you have the right to publish, remove sensitive information, and
assign someone to support the project afterwards. For those shared steps, follow
the release process and the release rules.
This page covers only what differs because the artifact is a model.
What differs from a code release
| Aspect | Source code | AI model |
|---|
| What you publish | Source code | Weight files and a model card |
| Licensing | One license for the code | Model license and training dataset licenses, judged separately |
| Documentation | README, contribution guide | Model card covering intended use, limits, bias, evaluation |
| Sensitive material | Code and commit history | Also personal data and copyrighted works inside the training data |
| Regulation | Export control (ECCN) | Also the documentation duties of the EU AI Act and Korea’s AI Framework Act |
Training datasets are where teams most often get stuck. Even with the model license settled, the
license of the data you trained on may restrict redistribution or commercial use. Check the two
separately.
Summary

- Obtain internal approval and request an OSRB review (stage A of the release process).
- Confirm the license of both the model and every training dataset. Do this early — if it fails
here, the rest of the preparation is wasted.
- Work through the pre-release checklist. Alongside it, write the
model card — the document your users will read first.
- Generate an AI SBOM to check your documentation against regulation.
- Publish the repository and operate it (stage D of the release process).
Stage D is written for source code going to GitHub. A model goes to your model development
organization’s Hugging Face account rather than a personal one; everything else about operating it
is the same. Ask your organization’s owner if you need access.
Checking your model before you publish
Work through the rights and data items on the pre-release checklist first. A private
repository is still an upload to an outside service, and pushing weights that turn out to contain
something you cannot publish is hard to undo.
After that, you can check the model yourself while the repository is still private. Push the
model privately and run BomLens, the SBOM generator, with your own Hugging Face token (HF_TOKEN);
it reports what is missing and how to fill it. Strengthen the model card with that result ahead of
time, and the OSRB review has the documentation it needs and goes more smoothly. The command to run
BomLens, how to prepare the token, and how to read the result are in AI SBOM.
Related pages
For questions and review requests about releasing an AI model, contact the OSRB
(opensource@sktelecom.com).
1 - AI Model Pre-Release Checklist
Check everything that has to be settled before a model goes public.
Read through this when you start writing the model card, not on the day you publish. Some items
here can only be fixed by training again.
1. The right to publish
Settle this first. If it fails here, the rest of the preparation is wasted.
Commonly missed
The model license and the training dataset licenses are separate. Settling on Apache-2.0 for the
weights does not help if a training dataset permits non-commercial use only. Check both.
2. Choosing a license
For the characteristics of each license, see
AI Model Licenses,
the Llama 2 guide and the RAIL guide.
3. Data and sensitive material
4. Model card
For how to write it, see Model card.
5. Files and identifiers
6. Approval and afterwards
Checking it automatically
The model card, license and dataset items above can be checked with a tool. Give BomLens a model
id and it reads the model card, then reports what is missing and how to supply it. Why it matters
and how to run it are both in AI SBOM.
What a tool can check stops at what the model card and the repository metadata reveal. Whether
personal data ended up in the training set, or whether a contract allows publication, is a
judgement a person has to make.
Related pages
For items you are unsure about, contact the OSRB (opensource@sktelecom.com).
2 - Writing the Model Card
What a model card is and which fields you need to fill in.
A model card is not a separate form. It is the README.md file in the model repository. The YAML
block at the top is the metadata machines read; the Markdown below it is the description people
read.
It is the first thing users read after release, and it is also where documentation for regulation
starts. Anything absent from it cannot be verified by any tool either.
File structure
---
license: apache-2.0
language:
- ko
- en
base_model: Qwen/Qwen2.5-7B
datasets:
- HuggingFaceFW/fineweb
pipeline_tag: text-generation
library_name: transformers
---
# Model name
Description, intended use, limitations and evaluation results.
The YAML block drives search and filtering, and it is what a tool reads when building an AI SBOM.
These four fields matter most.
| Field | Content | If left empty |
|---|
license | The license identifier for the weights | Users cannot tell on what terms they may use it |
datasets | Hub identifiers of the training datasets | Data provenance cannot be traced, leaving a gap for regulation |
base_model | The original model of a fine-tune, adapter (LoRA and the like), quantization or merge | The derivative relationship and inherited license duties stay hidden |
pipeline_tag | The task the model performs | The intended use is unclear |
For a non-standard license, set license: other and give the name and a link.
license: other
license_name: License name
license_link: https://example.com/license
Declare datasets as far as you can even when you are not publishing the data itself. If something
cannot be disclosed, saying so in the body — and why — is better than leaving it blank.
What belongs in the body
Start from the
model card template
that Hugging Face provides. These sections connect directly to documentation duties, so it is worth
not leaving them empty.
Model description
What the model does, how it is built, and roughly how many parameters it has.
Intended use
The situations the model was built for. This corresponds to the part of the EU AI Act’s technical
documentation that states a system’s purpose.
Out-of-scope use
Where it should not be used. Naming high-consequence domains explicitly — medical diagnosis, credit
scoring — is more useful than a general disclaimer.
Limitations and bias
The performance limits and biases you know about. If an area went unverified, write that it went
unverified. Disclosing a known problem is safer than omitting it.
Training data
What you trained on, and how you preprocessed and filtered it. If you filtered out personal data or
copyrighted works, say how.
Evaluation
What you measured, against what, and with what result. Name the evaluation datasets and metrics.
Where questions and reports should go. Someone has to be assigned to answer them after release.
When to write it
Leaving the model card until the end means reconstructing details from training time, which makes it
inaccurate. Recording as you go is far easier.
- Before training, record the dataset list, provenance and licenses.
- During training, record hyperparameters and evaluation results.
- While preparing the release, move these into the model card and add intended use and limitations.
- Confirm what is left with the pre-release checklist.
Related pages
For anything you are unsure about, contact the OSRB (opensource@sktelecom.com).
3 - AI SBOM and Regulation
Build an inventory of your model and check how far its documentation goes.
An AI SBOM is a machine-readable inventory of a model, its training datasets, and their provenance
and licenses.
- Where a software SBOM carries package dependencies, an AI SBOM also covers what a software SBOM
leaves out: model weights and training data.
- It is built from what the model card says, so a thin model card leaves an AI SBOM with that many
empty entries.
Why produce one
It serves two purposes.
- For the publisher it is a self-check on how complete the documentation is. Seeing what is missing
as a list means you can close the gaps before release.
- For the recipient it is evidence for a review. Just as SK Telecom asks suppliers for an SBOM,
other organizations have begun asking the same of AI models.
Key regulations
Not a compliance determination
Nothing on this page certifies or determines compliance with any regulation.
- It makes documentation gaps visible so a person can prepare.
- Interpreting them against a specific system’s legal obligations is a person’s job; when in doubt,
consult Legal and the OSRB.
| Regulation | Applies to | Applicable from | Key provisions |
|---|
| EU AI Act | Releasing a model that can serve general purposes | 2 August 2025 | Article 53 |
| EU AI Act | A model that ends up in a high-risk system | 2 December 2027 | Article 11, Annex IV |
| Korea’s AI Framework Act | Affecting the Korean market or its users | In effect since 22 January 2026 | Article 31 (transparency), 32 (safety), 33–34 (high-impact AI), 35 (impact assessment) |
- What reaches you when you release a general-purpose model is Article 53. It asks for technical
documentation, a copyright policy, and a public summary of training content, and it already
applies.
- Fine-tuning someone else’s base model may leave you outside those duties. The Commission’s
guidelines use a third of the original training compute as the dividing line.
- Releasing the weights and the architecture under an open-source licence exempts the technical
documentation, but the copyright policy and the
public summary of training content
remain. The summary follows a template the Commission provides.
- Article 11 and Annex IV attach to the high-risk system your model ends up in, not to the model.
- Korea’s AI Framework Act keeps its detailed documentation requirements in the enforcement decree,
so it is less specific than the EU’s. Article 32 (safety) applies only to systems trained above a
compute threshold set by that decree, so a linked element points at the subject of the duty rather
than establishing that the duty applies.
The G7 minimum elements
In May 2026, the cybersecurity agencies of the G7 jointly published “Software Bill of Materials for
AI — Minimum Elements”, with Germany’s BSI and Italy’s ACN leading the work. It defines 50 minimum
elements in seven clusters that an AI system’s inventory should carry. It is a non-binding
recommendation, not a regulation.
| Cluster | Elements | Needing human judgement | Content |
|---|
| Metadata | 10 | 0 | Who produced the inventory, when, with which tool |
| System-level properties | 9 | 4 | System name, application area, data flow |
| Models | 13 | 0 | Model identifier, license, integrity, training properties |
| Dataset properties | 10 | 5 | Dataset provenance, statistics, sensitivity, license |
| Infrastructure | 2 | 0 | Software dependencies and hardware |
| Security properties | 4 | 3 | Security controls, policy, vulnerability handling |
| Key performance indicators | 2 | 1 | Security metrics and operational performance |
| Total | 50 | 13 | |
- Thirteen of the 50 have no automated source. Things like the intended application area or the
sensitivity of the training data cannot be proven by any model card field, so a person has to
supply them.
- BomLens reports these 50 as 51 checks, because it scores model openness on its own line.
- The technical documentation the EU AI Act asks for corresponds to the G7 system-level, model, and
dataset clusters. That correspondence is BomLens’s reading.
- Advisory as they are, the elements overlap substantially with the documentation those regulations
require. Filling the G7 side also builds most of what a regulatory submission needs.
Producing an AI SBOM
Give BomLens, SK Telecom’s open-source SBOM generator, a
model id and it reads the model card, builds a CycloneDX AI SBOM, and reports G7 element coverage
alongside the regulatory mapping. (It reads the model card and the repository metadata; it does not
download the weight files themselves.)
- Setup, usage and how to read the reports are covered in the
BomLens AI model guide.
- It needs a Docker engine, and pulls the AI-model scanning image (about 3.5 GB) once.
- The command below assumes BomLens is already installed.
./scripts/scan-sbom.sh --project my-llm --version 1.0.0 \
--model "my-org/my-llm" --generate-only
A private repository needs a token with read access.
- Create a read-scoped token under Access Tokens in your Hugging Face account settings. A
fine-grained token needs read access to that repository granted explicitly.
- How to pass it as
HF_TOKEN, and what a gated repository additionally requires, are in
Private and gated models.
You get:
- an AI SBOM covering the model and its datasets
- per-element coverage, and how to supply what is missing
- the mapping to the EU AI Act and Korea’s AI Framework Act
- components whose license needs human review

Checking the result
An overall pass does not mean every G7 element is filled. G7 elements are all advisory and do not
move the verdict. Read the covered count and the gap count separately.
- Gaps come in two kinds: those you close by writing in the model card, and those a person has to
judge. Start with the first kind.
- To see what a report looks like before running anything, open the
conformance report
from a sample model scan. It opens with coverage per cluster and the licences flagged for review,
then works down to the per-check verdicts.
Related pages
References
For questions about producing or reading an AI SBOM, contact the OSRB
(opensource@sktelecom.com).