Image Registry Service
The Image Registry Service is a lightweight yet essential component of the Lakehousecat scaling infrastructure, responsible for hosting and distributing container images within the Kubernetes cluster.
Overview
The Image Registry Service serves as the central repository for all container images required by Lakehousecat and the underlying framework. During initial deployment, images are deployed to the registry service and subsequently distributed to individual services as needed.
Only administrators can modify scaling settings for the Image Registry Service. Access these settings through Admin Workspace > Settings > Services > Image Registry Service.
Primary Functions
- Image Hosting: Stores and manages container images for all Lakehousecat components
- Image Distribution: Provides images to individual services within the Kubernetes cluster
- Version Management: Maintains different versions of images for deployment flexibility
- Resource Optimization: Efficiently manages storage and bandwidth for image distribution
Default Configuration
The Image Registry Service is configured with minimal resource requirements due to its typically low scaling needs:
| Setting | Default Value |
|---|---|
| Autoscaling | Disabled |
| Replica Count | 1 |
| CPU Request | 100m |
| Memory Request | 128Mi |
| CPU Limit | 100m |
| Memory Limit | 128Mi |
Scaling Considerations
When Scaling is Needed
Based on extensive testing, scaling the Image Registry Service is rarely necessary under normal operating conditions. However, scaling may be beneficial in the following scenarios:
- High Concurrent Deployments: Multiple services pulling images simultaneously
- Large Image Sizes: Frequent distribution of large container images
- Network Constraints: Limited bandwidth requiring multiple registry instances
- Geographic Distribution: Multiple cluster locations requiring local registries
Scaling Options
The Image Registry Service supports both horizontal and vertical scaling approaches:
Horizontal Scaling
- Purpose: Distribute load across multiple registry instances
- Benefits: Improved concurrent access, better fault tolerance
- Considerations: Requires image synchronization between instances
Vertical Scaling
- Purpose: Increase resources for individual registry instances
- Benefits: Better performance for large images, improved response times
- Considerations: Single point of failure remains
Configuration Procedures
Accessing Scaling Settings
-
Navigate to Admin Interface
Admin Workspace → Settings → Services → Image Registry Service -
Verify Administrator Privileges
- Ensure you have administrator-level access
- Scaling modifications are restricted to administrators only
Enabling Autoscaling
To enable autoscaling for the Image Registry Service:
- Access the service configuration panel
- Toggle Autoscaling to Enabled
- Configure scaling parameters:
- Minimum Replicas: Recommended 1-2
- Maximum Replicas: Based on cluster capacity
- CPU Threshold: 70-80% for optimal performance
- Memory Threshold: 75-85% to prevent OOM errors
Manual Scaling Configuration
For environments where autoscaling is not desired:
- Set Replica Count to desired number of instances
- Adjust resource limits based on workload requirements:
- CPU Limits: Scale based on concurrent image pulls
- Memory Limits: Consider image cache requirements
Resource Planning
CPU Requirements
- Baseline: 100m CPU handles standard operations
- High Load: Consider 200-500m for environments with frequent deployments
- Burst Capacity: Ensure CPU limits allow for temporary spikes
Memory Requirements
- Baseline: 128Mi sufficient for basic registry operations
- Image Caching: Increase to 256-512Mi for better caching performance
- Large Images: Scale memory based on largest expected image size
Storage Considerations
- Image Storage: Plan for growth in image repository size
- Cleanup Policies: Implement image lifecycle management
- Backup Strategy: Ensure registry data is included in backup procedures
Best Practices
Performance Optimization
- Image Layering: Optimize Docker images for efficient layer sharing
- Registry Caching: Configure appropriate cache policies for frequently accessed images
- Network Placement: Consider network topology when placing registry instances
Security Considerations
- Access Control: Implement appropriate authentication for registry access
- Image Scanning: Integrate vulnerability scanning for stored images
- Network Policies: Restrict registry access to authorized services only
Monitoring and Maintenance
- Health Checks: Configure comprehensive health monitoring
- Log Analysis: Monitor registry logs for performance issues
- Capacity Planning: Regularly review storage and bandwidth usage
Troubleshooting
Common Issues
Image Pull Failures
- Symptoms: Services unable to retrieve images
- Solutions: Check registry connectivity, verify image availability
- Scaling Impact: Consider horizontal scaling for load distribution
Performance Degradation
- Symptoms: Slow image pull times, timeouts
- Solutions: Increase CPU/memory limits, enable autoscaling
- Monitoring: Check resource utilization metrics
Storage Issues
- Symptoms: Registry storage full, image push failures
- Solutions: Implement cleanup policies, expand storage capacity
- Prevention: Monitor storage growth patterns
Diagnostic Commands
For administrators troubleshooting registry issues:
# Check registry pod status
kubectl get pods -l app=image-registry
# View registry resource usage
kubectl top pods -l app=image-registry
# Examine registry logs
kubectl logs -l app=image-registry --tail=100
Integration with Other Services
The Image Registry Service integrates seamlessly with other Lakehousecat components:
- API Services: Provides base images for API containers
- Analytics Services: Supplies specialized analytics runtime images
- Data Processing: Distributes processing framework images
- User Interface: Serves frontend application containers
Performance Metrics
Monitor these key metrics to assess registry performance:
- Image Pull Success Rate: Should maintain >99% success rate
- Average Pull Time: Baseline measurement for performance degradation
- Concurrent Pull Capacity: Maximum simultaneous image retrievals
- Storage Growth Rate: Trend analysis for capacity planning
- Configure scaling parameters based on your specific use case requirements
- Test scaling configurations in non-production environments first
- Monitor resource utilization patterns before implementing permanent changes
- Consider implementing automated cleanup policies for unused images