# CuteNews Migration

`migrate_cutenews.php` is a one-time, read-only migration generator. It reads the active CuteNews installation and writes SQL for the internal CMS content system. It does not insert data into either database.

## Execution

The script is stored under `installation/migration/`, but it is intended to be copied into the application `src/` directory before execution. Its first requires are relative to `src/`:

```bash
cd src
php migrate_cutenews.php
```

An output path can be provided as the first argument:

```bash
php migrate_cutenews.php /path/to/cutenews-migration.sql
```

Without an argument, the SQL file is written next to the script as `migrate_cutenews.sql`.

The generated file must be reviewed and imported manually.

## Source Configuration

The script loads `CUTENEWS_CATEGORY_MAP` from the active `settings/config.php`. The map is the source of truth for which website sections are migrated and which CuteNews category belongs to each language.

Example:

```php
const CUTENEWS_CATEGORY_MAP = [
  'NEWS' => [
    'en' => 'News_English',
    'es' => 'News_Spanish',
  ],
];
```

The script does not assume that every section, language, or category exists on every installation. It also checks `AVAILABLE_LANGUAGES`; mappings for languages disabled in the current application are skipped and reported in SQL comments.

The following CMS categories are predefined and are not created by the migration. Their IDs are assigned directly to the category variables used by imported content:

| CuteNews section key | CMS category ID |
| --- | ---: |
| `NEWS` | `2` |
| `SERVER_DETAILS` | `4` |
| `FAQ` | `9` |
| `GUIDES` | `8` |

## CuteNews Reading Logic

The script loads CuteNews through its native `core/init.php` bootstrap. It then uses CuteNews functions rather than duplicating its storage format:

- `getoption('#category')` resolves configured category names to category IDs.
- `db_index_load('')` reads active article index entries.
- `db_index_load('archive')` reads archived article index entries.
- `db_get_nloc()` resolves an article timestamp to its storage file.
- `db_news_load()` loads the article payload.
- `cn_modify_title()` renders the legacy title.
- `cn_modify_full_story()` renders the legacy full body, including short-story concatenation and CuteNews formatting rules.

The active and archive lists are filtered by the category ID resolved from the configured category name. Articles are sorted by their Unix timestamp from oldest to newest.

## Translation Matching

Articles are not linked by their CuteNews IDs. The migration intentionally follows the agreed manual-review strategy:

1. Load every configured language list for a section.
2. Sort every list from oldest to newest.
3. Match article `1` in one language to article `1` in the other languages, then article `2`, and so on.
4. Create one CMS content node for each position that exists in at least one language.
5. Add only the translations available at that position.

This can produce incorrect pairings when publications were added or removed in only one language. The generated SQL includes a `-- REVIEW` comment whenever language counts differ, including the count per language.

## Generated CMS Structure

For each configured section with at least one resolved CuteNews category, the script generates:

- One `category` row in `cms_content_nodes` and one category translation per imported language, unless the section uses a predefined CMS category listed above.
- A reference to the existing predefined CMS category when applicable.
- One `content` node per oldest-to-newest matched publication.
- One `cms_content_translations` row for every available translation.

Categories use:

- `list_style = 'news'`
- `children_sort = 'publish_desc'`

Content nodes use:

- The section category as `parent_id`.
- `sort_order` based on the oldest-to-newest publication position.
- The first available translation timestamp as `created_at` and `modified_at`.
- The first available translation timestamp as node-level `publish_at`.
- `sticky = 0`.

Content translations use:

- Rendered CuteNews title and body.
- CuteNews upload URLs rewritten from any `http(s)://.../CuteNews/uploads/` path, including application prefixes and ports, to `assets/uploads/content/`.
- `is_draft = 0`.
- A normalized slug globally unique across all languages.

Content nodes use the first available CuteNews publication timestamp as `publish_at` for all translations.

New categories and content use `LAST_INSERT_ID()` SQL variables, while predefined categories use their existing CMS IDs directly, so translations correctly reference the intended node IDs even if the target CMS table is not empty.

## Review Comments

The generated SQL can contain comments for conditions requiring human inspection:

- `SKIPPED`: a configured language is unavailable or a CuteNews category name was not found.
- `REVIEW ... translation counts differ`: language lists have different article counts.
- `REVIEW ... title was empty`: a fallback title was generated.
- `REVIEW ... body was empty`: the imported body is empty.

Do not remove these comments until the corresponding content has been checked.

## Import Requirements

Before importing:

1. Confirm the CMS schema includes `created_at` and `children_sort` on `cms_content_nodes`.
2. Review every category and language skip comment.
3. Review every translation-count mismatch.
4. Inspect generated titles, bodies, slugs, and publication dates.
5. Import the SQL manually in the CMS database.

The output starts a transaction and ends with `COMMIT`. If manual review finds an issue, edit the generated SQL or roll back the import before committing.

## Deliberate Limitations

- The tool does not modify CuteNews files.
- The tool does not modify CMS tables directly.
- It does not infer translation relationships from titles, aliases, or content similarity.
- It does not migrate comments, CuteNews view counts, authors, or unrelated CuteNews metadata.
- It does not assume the category map shown in documentation matches the installation; the active configuration is always used.
