Key Value Service
The Key Value Service is a critical caching component that provides centralized data storage and retrieval capabilities for all Lakehousecat core services. This service acts as the primary caching layer, ensuring optimal performance across the entire platform.
Overview
The Key Value Service serves as the backbone for data caching operations, enabling fast access to frequently used data and reducing load on primary data sources. Its distributed architecture supports both small-scale deployments and enterprise-level operations with high availability requirements.
Only administrators can modify scaling settings for the Key Value Service. Access these settings through Admin Workspace > Settings > Services > Key Value Service.
Architecture Options
The Key Value Service supports two distinct architectural approaches to accommodate different deployment scenarios:
Stand-Alone Architecture
- Use Case: Small to medium deployments
- Benefits: Simple configuration, lower resource overhead
- Limitations: Single point of failure, limited horizontal scaling
Replication Architecture
- Use Case: Production environments, high-availability requirements
- Benefits: Data redundancy, improved fault tolerance, distributed load
- Considerations: Higher resource requirements, more complex configuration
Configuration Components
The service configuration is divided into two main areas:
Master Configuration
Controls the primary Key Value Service instance that coordinates data distribution and maintains cluster state.
Replica Configuration
Manages the replica instances that provide data redundancy and load distribution across the cluster.
Default Configuration
Master Configuration
| Setting | Default Value |
|---|---|
| CPU Request | 100m |
| Memory Request | 128Mi |
| CPU Limit | 500m |
| Memory Limit | 512Mi |
Replica Configuration
| Setting | Default Value |
|---|---|
| Replica Count | 3 |
| CPU Request | 100m |
| Memory Request | 128Mi |
| CPU Limit | 500m |
| Memory Limit | 512Mi |
Core Functionality
Caching Operations
- Data Storage: Temporary storage of frequently accessed data
- Cache Invalidation: Automatic cleanup of expired or outdated entries
- Memory Management: Efficient allocation and deallocation of cache space
- Performance Optimization: Fast read/write operations for core services
Service Integration
The Key Value Service provides caching support for:
- API Services: Session data, authentication tokens
- Data Management: Query results, metadata caching
- Analytics Services: Computed metrics, temporary calculations
- User Management: User preferences, role information
Scaling Considerations
When to Scale
Consider scaling the Key Value Service in the following scenarios:
Cache Hit Rate Degradation
- Indicator: Decreasing cache hit ratios
- Solution: Increase memory limits or replica count
- Monitoring: Track cache performance metrics
High Memory Utilization
- Indicator: Memory usage consistently above 80%
- Solution: Vertical scaling (increase memory limits)
- Prevention: Implement cache cleanup policies
Increased Load Distribution Needs
- Indicator: CPU utilization spikes on master instance
- Solution: Horizontal scaling (increase replica count)
- Benefits: Better load distribution across instances
High Availability Requirements
- Indicator: Business requirements for zero downtime
- Solution: Ensure minimum 3 replicas for fault tolerance
- Configuration: Enable automatic failover mechanisms
Scaling Timing
Always perform scaling operations outside of business hours to minimize impact on core services. Test scaling configurations thoroughly before implementing in production environments.
Recommended Timing:
- Maintenance Windows: During scheduled maintenance periods
- Off-Peak Hours: When system load is minimal
- Pre-Deployment: Before major application updates or releases
Configuration Procedures
Accessing Service Settings
-
Navigate to Admin Interface
Admin Workspace → Settings → Services → Key Value Service -
Verify Access Permissions
- Confirm administrator-level privileges
- Ensure service modification rights are enabled
Master Configuration Scaling
Vertical Scaling (Resource Adjustment)
CPU Scaling:
# Recommended CPU scaling increments
Light Load: 100m request, 500m limit
Medium Load: 200m request, 1000m limit
Heavy Load: 500m request, 2000m limit
Memory Scaling:
# Recommended memory scaling increments
Light Caching: 128Mi request, 512Mi limit
Medium Caching: 256Mi request, 1Gi limit
Heavy Caching: 512Mi request, 2Gi limit
Replica Configuration Scaling
Horizontal Scaling (Replica Adjustment)
Replica Count Guidelines:
- Minimum: 1 replica (development only)
- Standard: 3 replicas (production baseline)
- High Load: 5-7 replicas (peak performance)
- Enterprise: 7+ replicas (maximum availability)
Resource Scaling per Replica:
# Scale resources based on expected load per replica
Standard: 100m CPU, 128Mi Memory
Enhanced: 200m CPU, 256Mi Memory
Premium: 500m CPU, 512Mi Memory
Architecture Selection Guide
Choosing Stand-Alone vs. Replication
Stand-Alone Configuration
Suitable For:
- Development environments
- Small user bases (< 100 concurrent users)
- Non-critical applications
- Resource-constrained deployments
Configuration Example:
architecture: standalone
master:
replicas: 1
resources:
requests: { cpu: 100m, memory: 128Mi }
limits: { cpu: 500m, memory: 512Mi }
Replication Configuration
Suitable For:
- Production environments
- High availability requirements
- Large user bases (> 100 concurrent users)
- Mission-critical applications
Configuration Example:
architecture: replication
master:
replicas: 1
resources:
requests: { cpu: 200m, memory: 256Mi }
limits: { cpu: 1000m, memory: 1Gi }
replicas:
count: 3
resources:
requests: { cpu: 100m, memory: 128Mi }
limits: { cpu: 500m, memory: 512Mi }
Performance Optimization
Cache Strategy Configuration
Cache Policies
- TTL (Time To Live): Configure appropriate expiration times
- LRU (Least Recently Used): Enable automatic cleanup of old entries
- Size Limits: Set maximum cache sizes to prevent memory overflow
Memory Management
- Allocation Strategy: Configure memory pools for different data types
- Garbage Collection: Optimize cleanup intervals for performance
- Monitoring: Implement memory usage tracking and alerting
Network Optimization
Connection Pooling
- Client Connections: Optimize connection pool sizes
- Inter-Replica Communication: Configure efficient cluster communication
- Load Balancing: Distribute requests across available replicas
Monitoring and Metrics
Key Performance Indicators
Cache Performance
- Hit Rate: Percentage of successful cache retrievals
- Miss Rate: Frequency of cache misses requiring data source queries
- Response Time: Average time for cache operations
- Throughput: Number of operations per second
Resource Utilization
- CPU Usage: Monitor processing load across instances
- Memory Usage: Track memory consumption and available capacity
- Network I/O: Monitor data transfer rates
- Disk Usage: Track persistent storage utilization
Monitoring Commands
# Check Key Value Service pod status
kubectl get pods -l app=key-value-service
# Monitor resource usage
kubectl top pods -l app=key-value-service
# View service logs
kubectl logs -l app=key-value-service --tail=100
# Check cache statistics
kubectl exec -it <pod-name> -- redis-cli info stats
Troubleshooting
Common Issues
High Memory Usage
Symptoms:
- Memory utilization above 90%
- Out-of-memory (OOM) kills
- Performance degradation
Solutions:
- Increase memory limits for affected instances
- Implement more aggressive cache cleanup policies
- Consider horizontal scaling to distribute load
Prevention:
- Monitor memory trends regularly
- Set up automated alerts for high memory usage
- Implement proactive cache size management
Cache Miss Rate Increase
Symptoms:
- Decreased cache hit ratios
- Increased response times
- Higher load on primary data sources
Solutions:
- Analyze cache access patterns
- Adjust TTL settings for frequently accessed data
- Increase cache memory allocation
Replica Synchronization Issues
Symptoms:
- Inconsistent data across replicas
- Synchronization lag warnings
- Data integrity errors
Solutions:
- Check network connectivity between replicas
- Verify replication configuration settings
- Consider reducing replica count temporarily
Diagnostic Procedures
Cache Health Check
- Verify Service Status: Confirm all instances are running
- Check Connectivity: Test inter-replica communication
- Validate Data Integrity: Compare data across replicas
- Performance Analysis: Review response time metrics
Performance Debugging
- Resource Analysis: Check CPU and memory utilization
- Network Monitoring: Verify connection pool status
- Cache Statistics: Review hit/miss ratios and patterns
- Load Distribution: Ensure balanced load across replicas
Security Considerations
Access Control
- Authentication: Configure secure access for service connections
- Authorization: Implement role-based access to cached data
- Network Policies: Restrict network access to authorized services
- Encryption: Enable data encryption in transit and at rest
Data Protection
- Sensitive Data: Implement appropriate handling for sensitive cached data
- Retention Policies: Configure automatic cleanup of sensitive information
- Audit Logging: Enable comprehensive access logging
- Compliance: Ensure adherence to data protection regulations
Integration Guidelines
Service Dependencies
The Key Value Service integrates with multiple Lakehousecat components:
- Authentication Service: Caches user sessions and tokens
- Data Processing: Stores intermediate computation results
- API Gateway: Caches API response data and rate limiting information
- Analytics Engine: Temporary storage for metric calculations
Configuration Synchronization
- Environment Consistency: Maintain consistent configurations across environments
- Version Management: Track configuration changes and versions
- Rollback Procedures: Prepare rollback strategies for configuration changes