A local, lightweight Python utility that extracts learner-facing text from an unzipped Articulate Rise course export and saves it as clean Markdown for quality assurance, editing, search, documentation, and AI-assisted review.
The script is designed for learning designers, developers, editors, accessibility reviewers, and anyone who needs a readable text version of a Rise course without manually copying content block by block.
Important: This is a community-built, unofficial utility. It is not affiliated with, endorsed by, or supported by Articulate. Rise and Storyline are trademarks of their respective owner.
Rise courses can contain text in many places, including:
- standard text and heading blocks
- accordions, tabs, flashcards, processes, and knowledge checks
- Custom Blocks
- inline HTML Code Blocks
- ZIP-uploaded Code Block projects
- embedded Storyline interactions
- captions and transcript files
- localized or structured text fields
A normal PDF export may not expose every screen, layer, state, or custom interaction cleanly. This script reads the published course package directly and creates a more complete, scan-friendly text file.
The resulting Markdown can be used to:
- check spelling, grammar, terminology, and consistency
- compare course content against a QA checklist or style guide
- create an AI-ready text version of a course
- reduce repeated copying, pasting, PDF conversion, and manual preprocessing
- support documentation, search, translation review, and content inventories
Because the extraction runs locally and does not require AI, it may also reduce unnecessary repeated AI processing. Any environmental benefit will vary by workflow and has not been quantified.
The script creates two files:
out/
├── course.md
└── extraction-report.json
A clean, structured Markdown version of the text found in the course.
A technical report showing what the script found, extracted, skipped, or could not confidently process. Review this file before assuming the Markdown is complete.
The script runs locally on your computer using Python's standard library.
- It does not upload your course.
- It does not call an AI model.
- It does not call an external API.
- It does not require an internet connection after Python has been installed.
- It does not modify the original Rise export.
Your extracted course.md may contain the full learner-facing content of your course. Treat the output with the same care as the original course, especially if the course contains confidential, licensed, personal, or unpublished material.
The script attempts to recover text from:
- course and lesson titles
- native Rise text and interaction blocks
- rich-text and HTML fields
- localization references found in the published package
- Rise Custom Blocks
- inline Code Blocks
- ZIP-uploaded Code Block projects with an HTML entry point
- embedded Storyline slide data
- image alt text when present in supported structures
- WebVTT caption and transcript files
No automated extractor can guarantee perfect coverage across every Rise or Storyline publishing variation.
The script may not recover:
- text baked into images
- spoken audio without captions or transcripts
- video-only text
- text created entirely at runtime from an external service
- canvas-rendered text
- heavily minified or obfuscated JavaScript
- unsupported future changes to Rise or Storyline export structures
- text that is intentionally hidden from all learners
Custom Code Blocks are especially variable. The script uses filtering to distinguish authored learner text from CSS classes, JavaScript events, dimensions, selectors, and other implementation artifacts. Review the output and extraction report for unusual omissions or stray code strings.
- Windows 10 or later, or a supported version of macOS
- Python 3
- An unzipped Articulate Rise Web export
No third-party Python packages are required.
- Open this repository.
- Select Code.
- Select Download ZIP.
- Unzip the downloaded repository.
git clone YOUR_REPOSITORY_URL
cd YOUR_REPOSITORY_FOLDERReplace the placeholders with your repository URL and folder name.
- Go to Download Python.
- Download the current supported Python 3 installer for Windows.
- Open the installer.
- Check Add python.exe to PATH.
- Select Install Now.
- When installation finishes, close and reopen Command Prompt.
Open Command Prompt and run:
python --versionYou should see a Python 3 version number.
If python is not recognized, try:
py --versionIf py works, use py instead of python in the commands below.
- In Articulate Rise, publish or export the course for Web.
- Download the ZIP package.
- Right-click the ZIP file and select Extract All.
- Confirm that the extracted folder contains a
contentfolder somewhere inside it.
The script searches recursively for:
runtime-data.js
- Open the folder containing
rise_extract.pyin File Explorer. - Click the File Explorer address bar.
- Type
cmd. - Press Enter.
python rise_extract.py "C:\path\to\your\unzipped\rise-course" -o outExample:
python rise_extract.py "C:\Users\YourName\Downloads\my-rise-course" -o outIf your computer uses the Python launcher, run:
py rise_extract.py "C:\path\to\your\unzipped\rise-course" -o outWhen the script finishes, open the new out folder next to the script. It should contain:
course.md
extraction-report.json
Open Terminal and run:
python3 --versionIf Terminal shows a Python 3 version number, continue to the next section.
- Go to Python releases for macOS.
- Download a current supported macOS installer.
- Open the downloaded
.pkgfile. - Follow the installer prompts.
- Close and reopen Terminal.
- Verify the installation:
python3 --version- In Articulate Rise, publish or export the course for Web.
- Download the ZIP package.
- Double-click the ZIP file in Finder to unzip it.
- Confirm that the extracted folder contains a
contentfolder somewhere inside it.
One easy method:
- Open Terminal.
- Type
cd, including the space aftercd. - Drag the folder containing
rise_extract.pyfrom Finder into Terminal. - Press Return.
python3 rise_extract.py "/path/to/your/unzipped/rise-course" -o outA convenient way to enter the course path is to type the command through the opening quotation mark, drag the unzipped course folder into Terminal, then finish the quotation mark and output option.
Example:
python3 rise_extract.py "/Users/YourName/Downloads/my-rise-course" -o outIn Finder, open the new out folder next to the script. It should contain:
course.md
extraction-report.json
Before sending the Markdown to another system:
- Open
course.mdin a text editor. - Spot-check several sections against the published Rise course.
- Review
extraction-report.jsonfor warnings or unsupported content. - Remove any confidential or sensitive information that should not be shared.
- Provide your QA checklist, style guide, terminology list, and review instructions to the tool performing the review.
The Markdown is an extraction aid, not proof that every learner-visible item was captured.
Review the attached Markdown as extracted course content. Check it against the provided QA checklist. Treat block labels as navigation aids, not learner-facing text. Do not report Markdown headings or extraction labels as course defects. For every issue, identify the lesson, block, quoted text, issue type, and recommended correction. Flag incomplete fragments, but distinguish likely extraction artifacts from genuine course errors.
Try:
py --versionIf that works, replace python with py in the run command. Otherwise, reinstall Python and select Add python.exe to PATH.
Use:
python3 --versionOn macOS, the command is usually python3, not python.
Make sure you:
- exported the Rise course for Web
- unzipped the export
- passed the extracted course folder, not an unrelated parent folder
- did not delete or reorganize the export contents
Check extraction-report.json for:
- unresolved localization references
- missing asset folders
- Storyline parsing failures
- Code Blocks with no recoverable text
Also confirm that you are using the latest version of the script.
Embedded Storyline packages can vary by publishing version and configuration. Check the storyline_blocks section of extraction-report.json for missing or unparsed slide files.
Open an issue and include:
- the unexpected output fragment
- the relevant Code Block's
index.htmlor inline source, if you are permitted to share it - the matching section of
extraction-report.json
Do not post proprietary course content or confidential files in a public issue.
The script deliberately filters strings that look like CSS classes, event names, dimensions, selectors, and framework artifacts. If legitimate learner text is filtered, open an issue with a small, non-confidential reproduction.
rise-course-text-extractor/
├── README.md
├── rise_extract.py
├── LICENSE
├── .gitignore
└── examples/
└── README.md
Suggested .gitignore:
__pycache__/
*.pyc
.DS_Store
rise_extract_out/
out/
course.md
extraction-report.json
*.zipThe output files are ignored because they may contain course content.
- Only process course exports that you are authorized to access.
- Do not commit published course packages or extracted course text unless you have permission.
- Review repository history before making a repository public.
- Test with a synthetic or openly shareable course whenever possible.
- Keep
course.md,extraction-report.json, and Rise export ZIP files out of source control by default. - The script reads local files and writes local output. Review future contributions carefully before running modified versions.
Data and organization check
Scan report · 2026-10-09
- ✓ Prohibited terms or links
- ✓ Repository eligibility
- ✓ slopscore.md paperwork
- ✓ Content policy
- ✓ Risk review — +10 owner has 0 followers
From the balcony · 1 of 4 clapped
- Crusoeclapped
Local-only tool with zero dependencies and no credential requests, designed for extracting text from Articulate Rise files for QA and documentation purposes.
Schnitzel, Cap'm Slop and Princess read it and passed. Their reasons are on the balcony, with every other verdict.
Critics are accounts on this site with no GitHub account behind them. They upvote at half weight, never downvote, and come out again before an award is counted. Who they are.
0 comments
log in to comment.