sprited/kimodo

Kimodo text-to-motion: SOMA and G1 models, sequential prompts, motion constraints, multi-sample generation, GLB/BVH/NPZ and Mixamo FBX export, with animated previews. Unofficial NVIDIA Kimodo deployment.

Public
11 runs

Kimodo

Generate 3D skeletal motion from text, with optional motion constraints. Unofficial community deployment of NVIDIA Kimodo, hosted by Sprited.

Preview and downloads

The default output is a GLB animation file, with an MP4 motion preview. Select output_format to download BVH, NPZ, FBX, or all supported formats instead. Files are returned directly; ZIP packaging is optional via output_delivery=bundle.

The model playground animates the GLB. Replicate’s individual prediction page currently shows its first pose; use the MP4 there to review motion.

Disable preview for motion files only. Enable include_auxiliary to also receive a self-contained interactive HTML skeleton viewer, and metadata JSON. These auxiliary files are off by default to keep the output panel focused.

Choose output_format for GLB, NPZ, BVH, or FBX, or use all:

  • GLB contains the original skeleton hierarchy and animation. It includes animated joint spheres and bone rods by default. Disable skeleton_mesh for joint-animation-only export.
  • NPZ is the native Kimodo motion data.
  • BVH is available for SOMA models.
  • FBX requires your uploaded Mixamo-rigged character in custom_fbx and a SOMA model. sample_index, yaw_offset, and scale control the retargeted export.
  • A metadata JSON records the model, seed, skeleton, timing and generated filenames.

Motion controls

A period separates sequential motion segments: A person walks. A person waves. duration is seconds per segment, not total length. Optional segment_durations (for example [2, 3]) specifies each segment separately.

Generate 1–16 samples, using 0.5–30 seconds per segment and 10–500 diffusion steps. Larger requests take longer and may exceed available GPU memory. The default is one sample with 100 steps. Motion is finite, not guaranteed to loop.

constraints_json accepts the upstream Kimodo JSON format for root paths, full-body keyframes, hands, feet and other end-effectors. postprocess reduces foot skating; root_margin controls root correction. G1 skips postprocessing.

Available weights

Bundled: SOMA RP v1/v1.1, SOMA SEED v1/v1.1, G1 RP and G1 SEED. Default: Kimodo-SOMA-RP-v1.1. These use the original Llama-based LLM2Vec encoder. API callers do not need Hugging Face credentials.

The SMPL-X interface is retained for compatibility, but its separately gated weights are not available in this public deployment. Choose SOMA or G1.

Attribution and limitations

Built with Meta Llama 3. Kimodo source, NVIDIA model weights, Llama, LLM2Vec adapters, and the FBX dependencies retain their respective licenses and terms. See the Kimodo model card and ComfyUI-Kimodo, whose FBX retargeter is used here.

This renders skeleton previews, not a textured human or robot. Generated motion may miss parts of the prompt. Real character rigs need visual inspection after retargeting. Cold starts include loading the text encoder and motion model.

GLB includes a visible animated skeleton stand-in by default (skeleton_mesh=true). Disable it for joint-animation-only export; original joint names, hierarchy and animation tracks are preserved.

Model created
Model updated